
Unlock the Power of Turnkey HPC Interconnects
High Performance Computing (HPC) systems are essential for researchers and businesses that require processing power for data-intensive tasks. However, configuring and managing HPC systems can be daunting, especially for those who need specialized technical knowledge.
In this webinar, we will discuss the importance of interconnects in HPC systems and how they can impact the overall performance of your workloads. We will explore the benefits of turnkey HPC solutions, which offer pre-configured, fully integrated systems that can be deployed quickly and easily, allowing you to focus on your research or business needs.
By attending this webinar, you will better understand the role interconnects play in HPC systems and how turnkey solutions can help you accelerate your research or business workflows. We welcome your questions and look forward to seeing you there!
Webinar Synopsis:
-
Introductions
-
What are Interconnects
-
Types of Interconnects
-
What is a Turnkey HPC Interconnect
-
Bandwidth for Medium and Small Clusters
-
Optimal Topology
-
The Role of Software
-
Considering Costs
-
Cosmic Ray Interference
-
Building the Perfect System
-
Measuring Performance
-
State of the Art Interconnect
-
Where is it All Heading
Speakers:
-
Justin Burdine, Director of Solution Engineering, CIQ
-
Rose Stein, Sales Operations Administrator, CIQ
-
Gary Jung, HPC General Manager at LBNL and UC Berkeley
-
Jonathon Anderson, Senior HPC System Engineer, CIQ
-
Alan Sill, Managing Director, High Performance Computing Center at TTU
-
David DeBonis, Computer Scientist, CIQ
-
Matthew Dosanjh, Senior Member of Technical Staff, Sandia National Laboratories
Note: This transcript was created using speech recognition software. While it has been reviewed by human transcribers, it may contain errors.
Full Webinar Transcript:
Narrator:
Good morning, good afternoon, and good evening wherever you are. Thank you for joining. At CIQ, we're focused on powering the next generation of software infrastructure, leveraging the capabilities of cloud, hyperscale, and HPC. From research to the enterprise, our customers rely on us for the ultimate Rocky Linux, Warewulf, and Apptainer support escalation. We provide deep development capabilities and solutions, all delivered in the collaborative spirit of open source.
Justin Burdine:
Well, hey everyone. Welcome to another CIQ webinar. Thanks for joining us. Rose, I was all excited to, oh, first of all, I'm joined by my partner in crime, Rose Stein here, and we are excited to be taken over the kingdom. I guess Zane's on a plane, so he couldn't make it this round. And I'm sitting in for him. And, we'll see how it goes. I don't know what they're thinking, Rose.
Rose Stein:
I imagine that he's watching though, and we'll probably get a couple of side messages of what are you guys talking about? What are you doing?
Justin Burdine:
Exactly. Exactly. Well, we'll see. We'll see. We're looking forward to having back next week, but this week, what are we talking about? What are we looking to discuss?
Rose Stein:
This is actually interesting and I am really excited to hear different people's perspectives on this. So it's unlocking the power of Turnkey HPC interconnects to accelerate workloads.
Justin Burdine:
Okay. Well, please please tell me we have some smart people to talk about this because I am very new to this. This will be new stuff for me. I know what an interconnect is, but it's been 20 years since I've even thought about that from an HPC perspective. So who do we have on? Let's unleash the, we got Gary. Gary, fantastic.
Rose Stein:
Awesome. Hey Gary,
Justin Burdine:
Welcome Gary.
Gary Jung:
Hi. Hi.
Rose Stein:
Nice to see you.
Justin Burdine:
Oh, look at that. I knew we had other people in the wings. All right. Well, let's go ahead and intro. Gary, we'll go ahead and start out with you.
Introductions [6:49]
Gary Jung:
Hi. My name is Gary Jung. I run the institutional high performance computing for Lawrence Berkeley Laboratory, and I also manage the UC Berkeley institutional high performance computing program also.
Justin Burdine:
Awesome, awesome. Jonathan?
Jonathan Anderson:
My name's Jonathan Anderson. I'm a solutions architect here with CIQ and I have a background in academic high performance computing.
Justin Burdine:
All right, Alan?
Alan Sill:
Alan Sill. Like Gary, I wear two hats. I run the High Performance Computing Center here at Texas Tech University, and I also am one of several co-directors of a multi university industry University Cooperative Research Center in Cloud and Autonomic Computing with funding from the National Science Foundation.
Justin Burdine:
All right. Dave?
David DeBonis:
Hi, I'm David DeBonis. I'm a computer scientist over here at CIQ, with background in HPC embedded systems, mostly focused on system software.
Justin Burdine:
All right. And we got a new face here, Matthew.
Matthew Dosanjh:
Hi, I'm Matthew Jo, senior member of technical staff at San Diego National Laboratories. I worked on HPC middleware and research around that, particularly MPI, and open and smart nicks.
Justin Burdine:
Awesome. Awesome. Well, thank you guys for joining us. I really appreciate it. Hopefully this you'll be able to fill us in on all sorts of questions that we've got here about HPC interconnects. So, Rose got the first question?
Rose Stein:
You gotta break it down for me, though, and, maybe Matthew we'll start with you. So you are a new face, hi, it's nice to meet you. Thanks for being here. Can you break
down what are interconnects and then why is this important to the HPC system?
What are Interconnects [8:43]
Matthew Dosanjh:
So, interconnects are our form of networking in a sense. We run on trusted systems with very low latency needs for networks for scientific applications so that we can run hundreds of iterations across thousands of nodes a second. And so having the overhead of ethernet and that network stack doesn't really work for our use case. So there's been a long history in different companies making Intergen-X ranging from tray to the InfiniBand stack. And now we have a new generation coming out. But the idea of having a low latency networking solution for a trusted user base is the big thing.
Rose Stein:
Awesome. I appreciate that. Thank you. So what are some of the different types of interconnects? Gary, you want to jump in there?
Types of Interconnects [10:01]
Gary Jung:
Boy, there's the standard ethernet interconnect that Matthew was talking about that everybody uses to connect all the compute nodes. But usually when we refer to an interconnect or a fabric for a high performance computing system, then we're looking for something that sometimes could be more a performant than an ethernet connection. So the one that most people have converged on these days is InfiniBand, and there's different generations of that. But the current generation of that, that is currently on the market, is 400 gigabits per second. So that is quite a bit faster than most people would deploy with ethernet. And so, in addition to the higher performance, you would also get low latency. So, it'd be very low latency for tightly coupled jobs where there's communication between the compute nodes.
Justin Burdine:
Hey, Gary. What, what are those? What are the actual physical connections? Or is that all optical?
Gary Jung:
It can be, if they're within the rack, you can use passive copper cables and they use a QSFP connection. And if you go out of the rack, then you're going to be required to use optical connections. InfiniBand can go quite a distance. You can probably go to other areas within your campus with just optics.
Alan Sill:
At some sacrifice of latency. That's correct, yes. So, I think we understand that the fundamental purpose of interconnects is to provide parallelism. The recent talk by Torsten Hoefler at the HPC advisor council, I'll see if I can find a link, talked about the three types of parallelism, data parallelism, pipeline parallelism, and operator parallelism. And that's all pretty abstract. But the idea of a high speed interconnect is to allow more than one node to be used at once, really a simple fundamental concept. We have to understand that there've been a lot of complications lately. And certainly the field has always explored alternatives. We've seen people try to produce switch list fabrics, all optical interconnects. Rockport is popular in that area, work in that area.
There's always been alternative technologies. Cray has its Slingshot which is derived from ethernet. Omnipath is still on path to coming back. Cornellis has been pushing, but I think Gary's right, that InfiniBand is what people are familiar with in the sort of generic plain vanilla HPC cluster. Other complications have to do with ways that we're putting accelerators into the mix. So GPU's often are deployed in sets with dedicated interconnects that bypass the conventional ones. And, also we have to unveil a hidden secret, which is that many clusters, most clusters use the interconnect fabric, not just to allow computations to take place across multiple nodes, but to get high speed access to storage. So there's a hidden overloading of a function there that sometimes trips us up a little bit.
Gary Jung:
I wanted to just toss in this just for a comparison, but when interconnects... when I used to look at the top 500 list, one of the things that was interesting is that you could look at the different rankings of the clusters, and it would approximately, if you compare an InfiniBand cluster to an ethernet cluster, it took approximately twice the compute power of the ethernet cluster to come up with the same score on the HP lin pack as an InfiniBand cluster, or at that time, even a marinette cluster because of the latency and just the HPL depends on the parallelism that is supported by the low latency that Allen was mentioning.
David DeBonis:
Alan, you mentioned Torsten and I think he said during SE this year, Torsten of course is going to say this, but networks are the future of HPC. And it is because they're moving so much data around, and usually a lot of jobs, I would say, are mostly communication bound. Some of the aspects of RDMA, other sorts of technologies, the separate channels, help to make it more efficient at the fabric layer and make it more parallelizable.
Rose Stein:
So there's another thing, guys, that as we're talking about this topic, another word that is fun. Because every time I ask somebody they have a different, a different answer. But it's talking about Turnkey HPC, right? So unlocking the power of Turnkey HPC interconnects to accelerate workloads. What is, what does that mean to you?
What is a Turnkey HPC Interconnect [16:02]
Alan Sill:
I have to confess that that left me a bit baffled. I have never found a Turnkey interconnect. They're highly cantankerous herds of thoroughbreds if you want to make a very stretched analogy.
David DeBonis:
Can we say CIQ is your Turnkey?
Alan Sill:
Well, so you're welcome to come over. We have several air connect problems I'd be happy to point you towards. But, it's a dark art. Well, maybe that's too extreme. It's an art. And the interplay between driver versions for the fabrics, different generations of fabrics, fabric interfaces on the same part need the care and feeding of the cables themselves. People don't realize that there's firmware in those cables. There's a little computer in each end of one. So it's not just stringing cables together.
David DeBonis:
Scalability...
Gary Jung:
Oh, sorry,
David DeBonis:
Go ahead. No, go ahead, Gary.
Gary Jung:
I just gotta make a comment. We used to say that every cluster is serial number one.
David DeBonis:
I think that's what I was going to say. It's very dependent on your system and the needs of your system and the workloads that come through it. So it is customized, just like any system would be customized to, let's say, a video gamer that plays video games on their high performance system that is their workstation. Who knows? So, you have to worry about scalability, fault tolerance, a number of areas that cost will be very particular to an installation.
Alan Sill:
Cost is a big factor. So I'm at a university, we build modest scale clusters, and I'm willing to make some controversial statements. I actually never understood people who build clusters at our scale or even a little larger that don't max out the bandwidth on their fabrics these days to the limit of what they can afford. I'll give you an example that two years ago we put in the first university based AMD Rome system in the country, modeled after one that had just been built for the hawk supercomputer in Germany. And shortly thereafter, a lot of these started to pop up. And I was mystified by people putting in HDR 100 connects to dual 64 core nodes, taking giant leaps backwards in bandwidths per core. Anyway I could go on like this for some time, but Matthew hasn't jumped in yet. I wanted to see if I...
Matthew Dosanjh:
I was going to say that the one thing that also is difficult about Turnkey, in a sense, is that these numbers have been progressively getting more and more complex. And that's one of the things that we struggle with. You look at NVIDIA's Melanox or the new Cray Slingshot, and they're providing all this additional technology on the nets, and we're still trying to figure out from a software perspective, how do we utilize this? And that makes it, you add more and more components, the harder it is to have a true Turnkey solution.
David DeBonis:
And tuning those components is an art in itself, right. Not just the sizing and deciding, I'm going to have these particular pieces of hardware and this topology. It's even just getting them performance tuned.
Jonathan Anderson:
I think any exercise in creating a Turnkey HPC solution, and let alone just specifically the network, is going to be an exercise in tuning something for a specific use case. And there's one set of selection that's selecting the components that are usable, DPU's, and things like that, that might be useful in your application or in your suite of applications. But once you have a software stack and hardware stack that works together, and we can just set aside the firmware has bugs and the drivers have bugs, all of this needs to be maintained and kept up to date. Once you have all that, you still have questions, like Alan mentioned, the aspect of cost, maybe one deployment doesn't need dual 64 core CPUs and the bandwidth to feed that. And understanding what an application can actually take advantage of, and then how broad your network needs to be in order to fit that. We'll talk about topologies a little bit, I think, later. But not everyone needs a full bisection bandwidth, non-blocking fabric across all of their nodes, depending on how big their app, how far their application could scale across multiple nodes in multiple cores. Can't hear you, Alan.
Rose Stein:
Oh, Alan, you're on mute or something.
Alan Sill:
Hello? Can you hear me?
Jonathan Anderson:
Yes.
Bandwidth for Medium and Small Clusters [21:38]
Alan Sill:
Yes. I hear that statement a lot, and I think it's objectively true at a rational level, but it's also... I'm willing to argue not true at all from medium and small scale clusters. If you are going to go subdividing the bandwidth of your cluster up into trees and islands you're going to introduce scheduling complications that are especially severe in modest scale deployments. You're going to spend your time waiting for large jobs to schedule to fit into whatever islands of good connectivity you have. And the dual 64 core is old right? It's 96 now and the dual nodes are hard to build. The small half U nodes. I don't think we're going to see any dual nodes for a while. There will be full U units, but nonetheless you, you're not going to have one HDR 400 or you can put one HDR 400 interconnect into that. And you will just have the bandwidth that I have out of my dual 64 core HGR 200s. Right. It's an arms race that's taking place in your back plane. And then the earlier discussion, I forgot to mention that there are other things competing for that backplane bandwidth. Things like hyper-converged system... What do you call it? Composable computing, right? The PCIE networks and everything.
The weakness I often point out is that's yet another fabric, but those things have to talk to your motherboards and use up bandwidth on them too. So you have to watch the bandwidth per core ratio, the fabric pneuma assignments and a bunch of things. Now, this is the point you made about topology, Jonathan. I think you're right. But I've been building full bisection bandwidth, modest scale clusters for a long time, and I've never regretted it. When you get big, and I put a link in our internal chat that maybe someone could drop into the external chat, the talk that Torsten gave talks about what he... What I like about him is he always makes these declarative statements that you're left to puzzle out. The future is specified, he says, and then he has lots of slides about how you really need all the bandwidths you can get locally. So I'm not sure what it all means, but I will say that the topic you're raising, that of topology, is very important for large scale clusters. But I'm willing to advocate that all medium and small scale clusters, anything less than a few hundred nodes these days, given the core densities, you should spend money to get good connectivity. So, Gary, you raised a question of ethernet topology and ethernet's, of course, the basis of the Cray fabrics and the Cray fabrics have, I think, caught up in terms of bandwidth. Do you have Cray deployments that you can look at?
Gary Jung:
Not that I manage though, but
David DeBonis:
That brings what Alan said earlier too, is your local node and its interconnects, its pneumo nodes, all of the things that are separate memory domains but shared bus for at least the PCI bus. Those are the trade-offs and the trading points when you're on a small system. And probably the focal points of you really want to maximize your node, I think, more than even your interconnectivity. I'm not an expert on that, it feels like you want to do as much locally as possible to avoid that data movement as much as possible.
Alan Sill:
We can look at why dual socket nodes became so popular in HPC. It's almost the default. It's certainly not a bite against universal, but it was partly because of the cost of the interconnect. They were trying to get the most bang for the buck, building clusters out of nodes. And at that time, they thought they could just put more sockets to share the same interconnect and rest the infrastructure. Now, I think it's time to question that when you have 128 cores in a socket, maybe you need a dedicated card per socket. Maybe we should be looking for other ways to optimize cost. And that leads right back into Jonathan's question of topology.
David DeBonis:
I think that's a cue, Jonathan.
Justin Burdine:
So how do you choose an optimal topology? I mean, is there, are there steps you take to dissect the work you're working on? Or how does that work?
Optimal Topology [27:25]
Alan Sill:
So, again, I have to refer to Torsten's talk because I was just really impressed with it. It was only a few days ago that it came out. And there's certainly been other summaries in the past, but it talks about these three types of parallelism, data, pipeline, and operator. And it also talks about the topic that Gary raised, which is the interconnects of accelerators. And it depends on your scale and what you're trying to do. He introduces, at least for me, something new that probably has been around a while, something called a hamming mesh, which has, he says, many configurations. But there are already alternatives to fat trees. There's a popular one that's butterfly, and then there's various types of tourists interconnects.
To me, again, they make the most sense when you have a big enough cluster to make the difference. But for small things like a few dozen GPUs, you may already benefit. For example, if you've just got a few, you can use NVIDIA's envy link and get far faster. And it connects between just a few. So the question comes, how do you connect a few dozen to a few hundred? And I don't have enough experience with GPUs to answer that question. For CPUs, I think I need less than a few hundred. You should just do a factory.
David DeBonis:
It's interesting because Matt's in the center that I used to be a part of as well, and his father, I think it was his father that coined the phrase co-design. But the whole center's concept, and it was a computer science research institute at the time, or computational science. And the whole reason for it was to get both the application developers, the middleware developers, the people that were doing the linear algebra packages, and the people that were doing the system software and the people that were doing some of the underlying hardware or investigating hardware to understand their neighbors really well. And that was the way, for at least supercomputers, that's the way that you get a highly tuned system that can get onto the top 500 for a small system. And, maybe there is a more blanket Turnkey solution.
The Role of Software [29:56]
Alan Sill:
Yes. So, David, you've raised an interesting point in passing here, which is the role of software. We've managed to gloss over it up to now, because I think largely of the work done by a number of people to develop the standards that we rely on now so that we can ignore the complications, so that we have, InfiniBand verbs and MPI. That's a very complex topic that usually does and historically has been in different states for different types of interconnects, I should say it that way. And I think the community has done a great job of coming together and trying to hide that level of complexity from the average user. It's still there, but I think we don't spend a lot of our time thinking about the details of getting the hardware to communicate because of the work of the committees.
Considering Costs [31:01]
Rose Stein:
Can you guys actually expand a little bit more on what you mean by price? Is it literally the price of the hardware? Gary, you mentioned energy. You said it takes double the amount of energy. So when we're talking about how much things cost, what are we considering?
Gary Jung:
I was saying that it took twice as much compute on an ethernet cluster as compared to an InfiniBand cluster to achieve the same result on the lin pack test. But if you're talking about price, I guess usually when you're designing clusters, there's the compute budget, there's a storage budget, and then there's an interconnect budget. And usually the interconnect on the more parallel systems, the more parallelism that you're going to do on the system, the higher percentage of the budget should go to the interconnect. And so it can go anywhere from an interconnect with a blocking factor, say two to one, three to one, even four to one to what Alan called a full section bisection bandwidth, where essentially all the nodes on one half of the cluster could talk to all the nodes on the other half of the cluster without any kind of blocking. Yes. So that's where the cost comes in.
Matthew Dosanjh:
It's also probably worth noting that, when we're talking about networking and InfiniBand and stuff like that, the tables themselves are even more expensive than what you'd think of from a more traditional ethernet network, right? Where the trust is for the network and the compute resources is especially for large systems, which is where I spend my time doing research can be astronomical in a sense. And then you have to worry about power costs and everything like that just to operate the machine.
Alan Sill:
Correct. So let's go back to a point that Gary raised. That typically you'll make connections within a rack on copper, and then maybe between racks or certainly longer distances with optical fiber. Copper cable could cost a few hundred bucks. Optical cable can cost $1,200 depending on what it is, so when I say a modestly sized cluster is easier to build with full fat tree, full non-blocking folded cloth network. If you want to be a computer scientist, that's one of the reasons. The cables just cost you less. And there are fewer switches involved. The bigger the systems, the more... There's roughly a, I mean, basically identically a two to one ratio, but roughly when you fold in storage between the leaf switches and the course, which is the number of switches, grows, the cluster cost of switches is non-trivial. It can be tens of thousands of dollars. And back to another point that Gary raised, the ethernet, sort of pre-slingshot ethernet, classic ethernet can be used to build clusters. And the reason that people were willing to put up with a lower performance, more chattiness, more latency and so forth than ethernet, was that the switches are, and cables are so much cheaper. So that's why, especially a lot of the early Beowolf clusters are built on ethernet.
David DeBonis:
So one of the things too, on cost, I remember when I first started in the group over at Sandia, We were really thinking about fault tolerance and power. It was a really big thing. Resilience from a fault was important because we were seeing that systems were failing because the chip counts were getting up there to the point where we were scaling by the thousands of flops each whatever cycle. I don't want to get into the whole, when are we going to get exit scale thing, but that they were going through, it was always pushed out a couple of years. But it did seem like the fault rate of the actual devices was going to be so substantial that it wasn't, the system would never make progress because it would always have to fall back on a checkpoint to get back up to speed and then go forward. So, when you have more chips, when you have more devices, that means you'll have more failures. And so that's something to keep in mind too. The maintenance cost and the uptime or downtime cost of your system, no matter what the size really.
Matthew Dosanjh:
As you add nodes, your meantime to failure just shrinks drastically. And when you start looking at positioning of your cluster, right? We're at 6,000 feet here, you go up to someplace in the mountains like Los Alamos, it becomes actually significant compared to having something closer to Berkeley, which is closer to sea level.
Cosmic Ray Interference [36:44]
David DeBonis:
It's interesting because cosmic rays, believe it or not, really are a thing. Matter of fact, one of Matt's colleagues did his dissertation on just that. I think on a study with a comparison between Los Alamos and might have been some Sandia stuff, but I can't recall.
Justin Burdine:
So how does that impact, I'm that's really curious. Is that just adding more errors or what is it?
Matthew Dosanjh:
Flips
David DeBonis:
Bit flips. Yep. Wow.
Alan Sill:
And there is, of course, error correction and retry. But overall MPI is still remarkably
fragile. So, Jeffrey Fox did a lot of work on looking at the comparisons between, for example, map reduce and classic MPI. And looking for ways that the MPI standards could be made more robust and resilient to failure. But pretty much if you lose a node during a calculation, you lose the whole calculation no matter how many nodes are in it. Still, even in spite of this kind of work, I think we need more such work as we build excess scale systems with very large number of nodes to points that the raises become very important. And we still hear, though they've done a lot to tamp it down, we still hear a lot about the reliability of the largest of the scale deployments.
Matthew Dosanjh:
I will just hop in here for a second and say that this has been an ongoing topic in the MPI forum. We're finally at a point where errors aren't necessarily fatal, which we got to maybe two years ago or something like that. But even an out of memory error would just kill your MPI application which seems a little insane in retrospect...
David DeBonis:
I should think, I believe Matt is on the MPI committee, right?
Matthew Dosanjh:
I am the Sandia representative for the MPI Forum and in the open MPI developers community.
David DeBonis:
And I want to plug your conference. Oh, he's also chair, co-chair with another former colleague and a friend for Hot Interconnects this year, which is the IEEE standard, I think is the 29th year...?
Matthew Dosanjh:
I believe so. I'm the vice chair. The chair is Taylor Groves out at Lawrence Berkeley National Labs. Yep. Working in NERSC. And if anyone has papers they want to send over it or we run a free virtual conference. So, there's a lot of good talks and there's no cost to attend.
David DeBonis:
Cool.
Rose Stein:
Nice. Thank you.
Alan Sill:
So, I'll just say on behalf of the community, thank you. Because this is incredibly important work and it does require detailed understanding.
Rose Stein:
So, I have a question, guys, and I would like to hear from every single one of you, if you don't mind. And definitely it's going to change based on what the use case is. But say for example, you're going to build a cluster, let's say for a university like Allen's University, what would be the perfect system? How big would it be? What interconnects would you use?
Building the Perfect System [40:44]
David DeBonis:
I'll go first because I have no idea how to address that. Other than to say my focus has always been on using the power that you have available and performance of a large cluster that is not utilizing its power cap is wasting money. It's throwing money on the floor. So my research that I did was focusing on power efficiency. So how well did an application or a system use its power envelope? And how well could you reconfigure those into a lower power state or into a smaller power window or a longer execution time to help load balance a system. So you can do that on any system and you can make any system somewhat performant. I guess that's a subjective term, but I think that if we start to look at systems instead of how much money can we spend and how much fabric can we have, and how much the new fancy stuff that we have out there, how complex can we make it if we start to look at our workloads and how can we tune them better?
And how can we really get that understanding of how well an application or a workload is utilizing the compartment it's running in. That would be my approach, because I don't know how to answer it otherwise.
Alan Sill:
Rose, let me back up a little from your question then and define the terms, right? When you say university, let's just sort of describe what that environment is. Gary's familiar with this as well. You guys do have a lot of customers in this area. So university workloads are typically highly mixed in terms of they don't have just one dominant kind of application, or if they do, it's maybe due to special funding circumstances of a group. But you have a mix of workloads from very small core counts up to very large. And if you ask yourself, I should have prepared a graph, I have a graph where I show this, the distribution of jobs versus core count is highly peaked at the low end of small core count jobs.
But if you ask yourself at any given time, how many of my cores are occupied by small core count jobs versus large core count jobs after the schedule is done, its work of balancing things out, you find that it's sort of the 10 to a hundred and a hundred to a thousand bins that get the most populated. So yes, there are a lot of very small core count jobs, but each larger core count job takes more cores. So if you ask yourself, how many of my cores are taken up? It turns out that the medium skill jobs, and I think a lot of these environments, will dominate. So then you have to ask yourself, what can I do in designing my fabric storage CPU, to get the most productivity out. That's where the questions I raised earlier come into play.
So, I find with a we have about 750 nodes here and with cluster on that scale having, non-blocking to apologies just helps the scheduler do its job more efficiently and stuffing all the little jobs into the corners where it can between the big jobs. That's not the only way to do it. Look at bridges, it's a couple thousand nodes, I think. They have an oversubscribed fabric with non-blocking in, I think it's not quite per rack. Maybe it's a couple racks expanse. They do non-blocking per rack and then oversubscribe between racks. But that means your schedule has to fit those jobs, those bigger jobs into a rack, right? So, now I say rack, but the densities are such that you really are talking about a switch pretty soon.
And say, you have a 40 core 40 port switch with 20 ports going to worker nodes and 20 to upstream. Then, or maybe even oversubscribed, then 20 nodes at 256 cores a node is a pretty big job. So these are the kinds of things that you have to factor in. And I think if you phrase it that way in terms of workload, then you can get out of the area of just talking about a university and look at other similar workloads. This might apply to a small bio medics firm or to a number of different settings. When you get to the really large scale topologies, that's where the fancy fireworks come out, right?
Justin Burdine:
So, it sounds like really what's driving this is really understanding your workloads and understanding what you're really trying to solve. Not coming out saying, how much can I buy? Where can I, how can I build this out? It really is reverse engineering it. Okay. Interesting.
David DeBonis:
And especially with the rise of so much compute power in a GPU and the need for data movement, I think that changes the whole equation. And I think Alan was mentioning that earlier. It's a really important aspect. You have to understand your community and your workloads. What's the future of the hardware that is coming out next. Because some of these supercomputers what they do is they do a first initial install and then they plan out within a number of years, a couple of years or so an upgrade so that they don't just do that upfront cost every single time. They have a longer term path so that they can get reuse out of them. Knowing what hardware is going to do for the next generation is an important aspect of understanding the system that you're building today and the longevity of that system for future.
Matthew Dosanjh:
I think the one thing that I can add here is that it also really depends on your user base. I was running a panel at XMPI last year, and it was on how we are going to use MPI from GPUs? And we had a very impassioned member of the audience saying why does this matter? In industry when we're using HPC, we only program for the CPU because that's what's most cost effective. It takes too much developer time for a small company to port everything to the GPU and do all that other stuff. And you know that there's a cost balance here of life. How much time am I going to spend writing my code versus how much money am I going to spend on my cluster, like in four hours or whatever. And so then for a university, it's really, you want to tailor it to your expected workloads. And if they need something more than that, there's resources like NERSC and Tach that they can get or that your users can get allocations on.
David DeBonis:
That's a good point.
Alan Sill:
We've also been talking about this just in terms of on-premises clusters, but the same considerations play out in the cloud. And, and the thing that distinguishes them is that, of course, the cloud tries to make its money by selling on demand. And fundamentally that means that they have to deploy at a scale that's sometimes much larger but still be able to partition it up for individual workloads. So there's recently been a paper pointed out by Glen Lockwood, which talks about the InfiniBand structure of the RMA structure of Azure storage, for example. And it's really a tour de force of doing exactly the things you need to do to have a very large scale fabric deployment, but still allows you to do things like take care of security considerations for keeping an enclave of that available only to that user. I'll see if I can find a link to that.
Rose Stein:
So guys, how is it measured? How do you measure the performance of your HPC interconnect?
Measuring Performance [50:00]
Justin Burdine:
And actually, I'd love to add on to that if we could. Because Alan, you were talking about how there are HPC systems that are, are connected differently, right? And in rack we're doing at this speed. So I'm always curious, how would you know where or how would you figure out along with that? How do you monitor it? How would you know where their, where your bottleneck is? That's for everybody.
Matthew Dosanjh:
So we have a lot of different ways of looking at this that I'm aware of at least. There's lint level benchmarks. I know NERSC just collaborated on these large scale noisy networks to identify what we focus on at Sandia is proxy applications which are a small application that does something that looks like science and communicates like science and uses the ram like science, but is easier to understand and doesn't have the test the core function or performance.
David DeBonis:
Matter of fact, the mantiva mini apps are very much focused towards different aspects of computing, different computing workloads. So certain workloads could be compute
bound, memory bound, communication bound. They could be doing a lot of exchanges point to point, all of the different sorts of ways that you would... It's almost like they're different behaviors of scientific codes condensed down into a small kernel so that you can get these sorts of benchmarks. And some of the work when I was over at Sandia, I did some multi-generational studies and they have a beautiful set of test beds. There, a number of clusters are different, some are heterogeneous and some are homogeneous, they're outfitted with some of the more interesting newer generation stuff that's coming out. And we're able to see at least to a small scale, we'll do a scaling study still. To a small scale 4, 32, 96 nodes, how an application or one of these mini apps, the kernels perform as you start to scale. And usually you can see how an application is going to scale just by a few different doubling your size a few times.
Alan Sill:
So with some caveats that many people don't handle statistics of such benchmarking properly, and that's another whole webinar we could do. But the standard thing to do is to study your scaling and practical terms. Let me point out, if you're going to ask for allocation on a national scale for computer, you're going to need to demonstrate that your application is able to scale. Unfortunately, I have a little constraint on what we're to say. We've actually doing some work within my center with one of the member companies on improving the state of the art for benchmarking of MPI code. But I mean, just go back to what I started with, it's easy to fool yourself with statistics when you're benchmarking. You need to be careful, more careful than most people are when they're doing these scaling studies.
Ultimately, what the scientist cares about is how many units of X workload they can get through. But even simple tools like total view or debuggers can help you look at how you're using your memory, how you're using dial. One of the things we have published from my center, and we've talked about in previous webinars, is that we have open source tools Github.com/nsfcac, that you can use on pretty much any modern system that supports Redfish, to gather out of band statistics on things like memory, power, CPU power. And we're working to fold in these bandwidth statistics so that you can do a touchless overview benchmarking of your cluster how, where is it spending time and effort without having an instrument code.
David DeBonis:
I think Redfish actually utilized the power api, which was a Sandia product that I worked on when I was there. And it's very important to get those out band information and not interfere with the workload or else wait, you're perturbing the system. But it goes back to cost, like you were saying earlier, Rose, asking about cost, it isn't really dollar cost, it's science. Like Alan had alluded to, it's the cost in science. How much science can I get done on this system and how do I get it done very efficiently. So, I can get more science done. It's as simple as that, I think.
Alan Sill:
So Gary, how often do you get researchers coming to you and saying, can you speed up my code versus just, I can't get this to run?
Gary Jung:
It does happen. Usually, most of the issues are these days that I can recall, like saying the last couple of months have been IO related. And so it hasn't been with the interconnect as more with the IO but it does. They do, they do.
Justin Burdine:
So, I guess we're getting towards the end. One of the things that I was curious about is what, because I talked to a bunch of customers who are looking for direction and I was curious what is the state of the art right now if you were to recommend an interconnect? And I know there's obviously depends on the workloads, but what's the state of the art? And Alan, I know you may not be able to say and dive into what you're talking about, but what's, if we were to go off the shelf, what would that be?
State of the Art Interconnect [56:31]
Alan Sill:
Listen, I think it's more competitive than people realize. Like Gary said, most people's starting point is InfiniBand. But let's face it, InfiniBand is expensive and touchy. I've got both an InfiniBand and an Omnipath deployment, and the Omnipath never gives me any trouble. The InfiniBand, I have to be careful with every version update, and there seem to be lots of ways that small communication areas can interrupt performance. I'd mentioned the switch list fabrics the Rockport style ones, you just put a different PCIE card. You don't have either InfiniBand or Omnipath, but it is limited in the number of nodes you can connect. So the great advantage of that is it's very little latency. I mean, essentially, no latency. Well, not none, but instead of PCIE back plan latency at that point.
Then we also have the PCI interconnects, the composable computing fabrics. So you can buy a box of GPUs and shove them in. And depending on your workload, again, if you need those GPUs to communicate or not communicate, you could save a lot of money if you don't want to spend a lot of money on interconnects by just having a box of GPUs. Running them separately. So all of the above. Okay. But I think most people do start out, as Gary said, looking at InfiniBand fabrics and that's a good place to get your feet wet.
Matthew Dosanjh:
I think, what was it, with Melanox now being owned by Nvidia and then Omnipath is spun out of Intel. There's a hardware dependency that is a little bit baked in to some of these.
David DeBonis:
I would say Affinity. I worried about the same sort of thing when Xylinks and Alterra left the marketplace and got consumed by MD and Intel, same sort of thing. They're smart, they know that's going to be the optimization path and they also know that the ecosystem and the people that are out there that can do those things are very limited. Even though Xylinks was trying to create the tools for an ecosystem of design engineers to make their own fabrics. So in a way you're steered, you're always a slave to the industry and the movement in the industry. So your decisions of what you get could have to do with who owns what, which is a sad reality, I guess.
Alan Sill:
So, we're running out of things that Nvidia hasn't bought, right?
David DeBonis:
Right.
Justin Burdine:
Well, going on the path, do we see GPUs as being where this is all heading? Or is there still a need for CPU based compute just for pricing or cost?
Where is it All Heading [59:49]
David DeBonis:
I think it depends on the workload. And yes, we definitely do see a need for that with the
amount of small compute and large data movement that you see in some of the machine learning codes and training.
Matthew Dosanjh:
Well, and, and it also depends on your users, right? I alluded to this earlier. We have a
bunch of people working on GPU codes, but there's a lot of people that don't want to deal with that. I don't want to deal with it. Added complexity to their science of having both a CPU show to manage everything and then moving stuff over to the GPU and moving it back, doing your communication. Or trying to figure out how to do MPI from the GPU. But that's in its early stages and has not been standardized. And we're working on that in the MPI forum, but it's not a reality yet.
Alan Sill:
I think it's a little informed by the topologies of the big clusters too, right? We've seen well, let's see. Frontier has four GPUs per CPU, if I recall correctly. Then I think the new AMD CPU plus GPU systems are just being deployed where you have essentially all on a socket or all within a small portion of the backplane. That's great. So my controversial take that I can never get hardware vendors to agree to, is that when they build a national scale, really big system, they should sell little units of that system to a bunch of universities. Just identical, because then we can get the most out of the software development. But I can never get any of the hardware vendors to agree that because they lose their shirts on these big systems.
Justin Burdine:
Well, we are coming up on the top of the hour. Any final thoughts from anybody?
David DeBonis:
Well, I want to thank Matt for doing this at the last minute. I just called him an hour before this to see if he can make it. I'm glad he could. Matt's a good friend of mine here in Albuquerque and it was good to see him
Justin Burdine:
Absolutely.
Matthew Dosanjh:
Thanks for having me.
Justin Burdine:
Sure. Love to have you again. This is great.
Rose Stein:
Thanks Matt.
Justin Burdine:
Any other final thoughts? Everybody else? Good? We talked it out.
Rose Stein:
Talked it out. Well you guys, make sure that you like and subscribe. We are here every single week, same time, same place, same channel. We will be here If you guys also have suggestions on topics for webinars, things that you want to dive into, please leave those in the comments. We want to read them and we want to serve our community the best that we can. So thank you everyone on the panel for being here. Thank you for watching. You are very amazing. And we'll see you next week, same time, see you guys.
Transcript
Good morning, good afternoon, and good evening wherever you are. Thank you for joining. At CIQ, we're focused on powering the next generation of software infrastructure leveraging the capabilities of cloud, hypers scale, and HPC. From research to the enterprise, our customers rely on us for the ultimate rocky Linux, werewolf, and appainer support escalation. We provide deep development capabilities and solutions, all delivered in the collaborative spirit of open source. Well, hey everyone. Welcome to another CIQ webinar. Thanks for joining us. Rose, I was all excited to uh I'm Oh, first of all, I'm joined by my partner in crime, Rose Stein here, and we are excited to uh to be uh taking over the kingdom.
I guess Zayn's on a plane, so he couldn't make it this round, and I'm sitting in for him. And uh we'll see. We'll see how it goes. I don't know what they're thinking. I I you know I imagine that he's watching though and we'll probably get a couple of like side messages of like what are you guys talking are you doing exactly exactly we'll see we'll see we're looking forward to have back next week but this week what are we talking about what are we looking to discuss you know this is actually interesting and I am really excited to hear different people's perspective on this so it's unlocking the power of turnkey HPC interconnects to excel accelerate workloads.
Okay. Well, please yeah, please tell me we have some smart people to talk about this because I am very new to this. This will be uh new stuff for me. I know what an interconnect is, but it's been 20 years since I've even thought about that in an HPC perspective. So, who do we have on? Let's uh let's unleash the Gary. Fantastic. Awesome. Hey, welcome Gary. Hi. Hi. Nice to see you. Oh, look at I knew we had other people in the wings. All right. Well, let's let's go ahead and intro. Gary, we'll go ahead and start out with you. Hi, my name is Gary John.
I run the institutional high performance computing for Lawrence Berkeley Laboratory and I also manage the UC Berkeley uh institutional high performance computing program also. Awesome. Awesome. Jonathan. Yeah, my name is Jonathan Anderson. I'm a solutions architect here with CIQ and I have a background in academic high performance computing. All right, Alan. Yeah, Alan Sil and like Gary, I wear two hats. I I run the high performance computing center here at Texas Tech University and I also am one of several co-directors of a multi-university university industry university cooperative research center in cloud and autonomic computing with funding from the National Science Foundation. All right. Dave. Yeah. Hi, I'm David Dabonis.
View full transcriptHide full transcript
Uh, I'm a computer scientist over here at CIQ. Um, with a background in HPC embedded systems, mostly focused on system software. Um, and then Yep. All right. And we got a new face here, Matthew. Hi, Matthew Dosa, senior member of technical staff at San Die National Laboratories. I work on uh HPC middleware um and research around that uh particularly MPI and uh open and smart nicks. Awesome. Awesome. Well, thank you guys for joining us. I really appreciate it. Hopefully this uh you'll be able to fill us in on all sorts of questions that we've got here about uh HPC interconnects. Yeah. So, you got the first question.
Yeah. I I I you got to break it down for me though and like maybe Matthew will start with you. So you are a new face. Hi, it's nice to meet you. Thanks for being here. Um can you break down like what are interconnects and then why is this important to the HPC system? So interconnects are kind of uh our form of networking. Um in a sense uh we run on trusted systems with uh very low latency needs for networks um for scientific applications so that we can run you know hundreds of iterations across thousands of nodes a second. Um and so having the overhead of Ethernet and uh that network stack doesn't really work for our use case.
So there's been a long history in different companies making uh interconnects ranging from you know Trey to uh the Infiniban stack and now um we have kind of a new generation coming out now. Um, but the idea of having a low latency networking solution for a trusted user base is kind of the big thing. Awesome. I appreciate that. Thank you. Um, so what are some of the different types of interconnects? Gary, you want to jump in there? Boy, you I there's uh the standard Ethernet interconnect that um Matthew was talking about that connects everybody uses to connect all the compute nodes. Um but usually when we refer to interconnect or fabric for a high performance computing uh system then we're looking for something that uh sometimes could be more performant than an Ethernet connection.
So the one that most people have converged on these days is Infiniband and there's different generations of that but the current generation of that uh that is currently on the market is um is 400 gigabits per second. So that is quite a bit faster than you most people would deploy with Ethernet. And so uh in addition to the higher performance you would also get low latency. So it'd be very low latency for tightly coupled jobs where there's communication between the compute nodes. And Gary, what what are those what are the actual physical connections or is that all optical? Um it can be uh if they're within the rack, you can use passive copper cables and they they use a QSFP connection.
Um and uh if you go out of the rack then you're then then you're going to be required to use optical connections. Okay. So you're going these days can go quite a quite a quite a distance. You can uh you can go probably to other areas within your campus uh with just optics at some sacrifice of latency. That's correct. Yes. Yeah. Um so I think you know we have to understand that the fundamental purpose of um interconnects is to provide parallelism. Um the um recent talk by Torston Hler of uh uh at the HPC advisor come out uh council I I'll see if I can find a link uh talked about the three types of parallelism data parallelism pipeline parallelism and operator parallelism.
And that's all pretty abstract, but the idea of a high-speed interconnect is to allow more than one node to be used at once. Really kind of a simple fundamental concept. Um, we have to understand that there have been a lot of complications lately. Um, and uh certainly the field has always explored alternatives. We've seen uh people try to uh produce um switchless fabrics uh all optical interconnection. Rockport is up popular in that area working in that area. There there are uh there's always been alternative um technologies. Craig has its slingshot um which is derived from Ethernet. Um Omnipath is is uh still on path to coming back.
uh Cornellis has been pushing but I think Gary's right that that Infiniband is is what um people are familiar with and in a sort of generic plain vanilla HBC cluster. Um other complications have to do with uh ways that we're putting accelerators into the mix. So GPUs often um are deployed in sets with dedicated uh interconnects that bypass the conventional ones. And and also we have to un you know unveil a hidden secret which is that many clusters most clusters use the interconnect fabric not just to allow computations to take place across multiple nodes but to get high speeded access to storage. So there's a kind of a hidden overloading of the function there that that sometimes trips us up a little bit.
Mhm. H um you know I I wanted to just just toss in this just for a comparison but I um when interconnects um when I used to look at the top 500 list um one of the things that was interesting is that you could look at the different rankings of the clusters and it would approximately for uh an if you compare an Infiniban cluster to an Ethernet cluster it took took approximately twice the compute power of the Ethernet cluster to come up with the same um score on the HPL lin uh as an infiniband cluster or at that time even a mirroret cluster with the because of the latency uh and this and and just the HPL is you know it it's a depends on the parallelism that is that is supported by the low latency uh that Helen was mentioning.
Yeah. Yeah. Alan, you mentioned Torstston and and I think he said during SC this year um Torstston of course is going to say this, but um networks are the the future of HPC. Um and it is because they're moving so much data around and usually a lot of jobs I would say are mostly uh communicationbound. um some of the aspects of uh of like RDMA um other sorts of technologies the separate channels those help to make it more efficient at the at the fabric layer and make it more paralyzable. So, so there's another thing guys that as we're talking about this topic, another word that is uh kind of fun because every time I ask somebody, they have a different a different answer, but it's talking about turnkey HPC, right?
So, unlocking the power of turnkey HPC interconnects to accelerate workloads. What is what does that mean to you? I have to confess that that left me a bit baffled. I I have never found a turnkey interconnect. They are they are highly canankerous herds of thorbreds if you want to make a very stretched analogy. Can we say CIQ is your turnkey? Well, so uh you know uh you're welcome to come over. We have several um air connect problems I' I'd be happy to um to point you towards. Uh but you know uh it's a dark art. Um well maybe that's too extreme. It's it's an art. Uh and you know the the the interplay between you know driver versions uh for the fabrics uh different generations of fabrics uh fabric interfaces on the same uh the care and feeding of of uh you know the cables themselves.
So, and you know, people don't realize that there's firmware in those cables and you know, there's there's a little computer in each uh end of one. So, you know, it's it's um it's not just stringing, you know, cables together. Yeah. Yeah. Scalability to say that Oh, sorry. Go ahead. No, go ahead, Gary. I was just going to make a comment. We used to say that every cluster is serial number one. Yeah, I think that's that's what I was going to say. It's very dependent on your system and the needs of your system and the workloads that come through it. So it is it is customized just like any system would be customized to let's say a video gamer that plays video games on their high performance system that is their workstation.
Who knows? Um so it is it is you have to worry about scalability, fall tolerance, a number of areas that will be very particular to an installation. Yeah, cost is a big factor. Um so I'm at a university we build you know modest scale clusters and um I've I'm willing to make some controversial statements. Uh I actually never understood people who build clusters at our scale or even a little larger that don't you know max out the bandwidth on their fabrics these days to the limit of what they can afford.
Uh I'll give you an example that uh two years ago we put in the first university based AMD Rome system in the country modeled after one that would had just been built the Hawk supercomputer in uh in um Germany and uh shortly thereafter a lot of these started to pop up and I was mystified by people putting in HDR 100 connects to dual 64 core nodes. uh you know taking giant leaps backwards in uh Bandwood's percore. Anyway, uh I could go on like this for some time, but uh Matthew hasn't jumped in yet. I wanted to see if any I was going to say that the uh um the one thing that uh also is difficult about turnkey in a sense is that these networks have been uh progressively getting more and more complex.
And that's one of the things that we struggle with is like uh you know you look at uh Nvidia's Melanox uh or uh the new tray slingshot and they're providing all this additional technology on the uh ny and we're still trying to figure out from a software perspective how do we utilize this and that makes it um you know as you add more and more components the harder it is to have a true turnkey solution. Yeah. Yeah. And tuning those components is a is is an art in itself, right? And not just the sizing and deciding I'm going to have these particular pieces of hardware and this topology.
It's it's even just getting them yeah performance tuned. I think any exercise in creating a turnkey HPC solution and let alone just specifically the network is going to be an exercise in tuning something for a specific use case. And there's there's one set of selection that's selecting like components that are usable like like DPUs and things like that um that that might be useful in your application or or in your suite of applications. Um but once you have a software stack and hardware stack uh that works together and we can just set aside the the firmware has bugs and the the drivers have bugs and all of this needs to be maintained and and kept up to date.
Once you have all that, you still have questions like Alan mentioned the the aspect of cost. Like maybe one uh one deployment doesn't need dual 64 core CPUs and the bandwidth to feed that. Uh and understanding what an application can actually take advantage of and then how broad your network needs to be in order to fit that. We'll talk about topologies a little bit I think later. Um, but not everyone needs a full bisection bandwidth non-blocking fabric across all of their nodes depending on how big their how far their application could scale across multiple nodes in multiple cores. Yeah. Can't hear you, Alan. Oh, yeah. Allan, you're on mute or something.
Hello. Can you hear me? Yes. Yes. Ah, yes. So I I I hear that statement a lot and I think it's objectively true at a rational level, but it's uh also I'm I'm willing to argue not true at all for medium and small scale clusters. Yeah. Uh if you're going to go um subdividing the bandwidth of your cluster up into uh you know trees and islands, um you're going to introduce scheduling complications that are especially severe in modest scale deployments.
uh you're going to spend your time waiting for large jobs to schedule to fit into whatever islands of good connectivity you have and uh you know the dual 64 core is you know old right uh it's uh you know 96 now and dual and you know the dual nodes are hard to build the small one you know halfu nodes you know I don't think we're going to see any dual nodes for a while there will be full U units but nonetheless uh you're not going to have one HDR400 um or you can put one HDR400 interconnect into that uh and you will just have the
bandwidth that I have out of my dual 64 core uh HDR200's right uh it's an arms race that's taking place in your back plane and in the earlier discussion I forgot to mention that there are other things competing for that uh backplane bandwidth uh the uh you know things like you know hypercon converge system converge what do you call it composable computing right the uh the PCIe uh networks and everything uh the weakness I often point out is it's yet another fabric but you know those things have to have to talk to your um your motherboards and use up bandwidth on them too so it's
you have to watch the the bandwidth per core ratio the fabric uh numa uh assignments and a bunch of Now, as to the point you made about uh topology, Jonathan, I think you're right, but I've been building full bisection bandwidth uh monoscale clusters for a long time and I've never regretted it. Uh uh when you get big and I I put a link in the in our internal chat that maybe someone could drop into the external um chat.
uh the um uh uh the talk that Torsson gave talks about uh what he he what I like about him is he always makes these declarative statements that you're left to puzzle out the future is sparsified he says and then and then he then has lots of slides about how you really need all the bandwidth you can get locally so I'm not sure what it all means but uh I will say that the topic you're raising that of topology is very important for large scale clusters but I'm willing to advocate that all medium and small scale clusters anything less than a few hundred nodes um
these days given the core densities you should spend you know money to get good connectivity so Gary you raised the question of Ethernet topology and Ethernet's of course the basis of the Cray fabrics and the cray fabrics have I think caught up in terms of bandwidth do you do you have cray deployments that you you look at. Not not that I not that I manage though, but um nurse get our site has has a has a slingshot. Yeah. And and you're right. I I was going to also add to what you were saying. I we're just starting to see the emergence of these um you know multi-GPU nodes and for the people are training large models then the communications is a is a is a bottleneck and so even a single 400 gigabit connection sounds a little light for an 8-way GPU system if you're doing uh training.
Yeah. Yeah. That's that kind of brings what Alan said earlier too is your your local node and its interconnects. it's it's numa nodes all of the things that are separate kind of memory domains uh but shared bus uh for at least the PCI bus those are the those are the tradeoffs and the and the trading points when you're on a small system and probably the focal points of you really want to maximize your node I think more than even your interconnectivity I'm not an expert on that but um I it feels like you want to do as much locally as possible to avoid that data movement as much as possible.
Yeah. So we can look uh you know at why dual socket uh nodes became so popular in HPC uh almost the default and certainly not binding universal but uh it was partly because of the cost of the interconnect. they were trying to get the most uh bang for the buck building um uh clusters out of nodes and that and it at that time they thought they could just put more sockets to share the same interconnect uh and the rest of the infrastructure. Now I think it's time to question that when you have 128 cores in a socket maybe you need a dedicated uh card per socket.
um maybe we should be looking for other ways to to optimize cost and that leads right back into Jonathan's question of topology. I think that's acute Jonathan. Yeah. So h how do you choose like an optimal topology like I mean is there are there steps you take to kind of dissect the work you're working on or how does that work? So again, I have to refer to Torston's talk because I I you know, I was just really impressed with it. It was only a few days ago that it came out and uh and there's certainly been other summaries in the past, but uh it um it it talks about these three types of parallelism, data, pipeline, and operator.
And it also talks about the topic that that Gary raised, which is, you know, the interconnects of of accelerators. Um and it depends on your scale and what you're trying to do. Um he introduces uh at least for me something new that probably has been around a while something called a Hamming mesh which is uh has he says many configurations. Uh but there are already alternatives to fat trees. There's you know a popular one is a butterfly and then there's a various types of Taurus interconnects.
uh to me again they they make the most sense uh when you have a big enough cluster to make the difference but for small things like uh a few dozen GPUs you may already benefit uh you know for example you just got a few you can use Nvidia's NVL link uh and get fire faster and it connects between just a few so the question comes how do you connect you know a few dozen to a few hundred um and uh I don't have enough experience with GPUs to to answer that question. Uh for CPUs, I think any less than a few hundred, you should just do a pad tree.
Yeah, it's interesting because one of the uh Matt's in the center that I I used to be a part of as well and his father, I think it was his father that coined the phrase codees. Um but um he the the whole c center's concept and it was the computer science research institute at the time or computational science um and the whole the whole reason for it was to get both the application developers, the middleware developers, the people that were doing the the linear algebra packages and uh the people that were doing the um the system software and the people that were doing some of
the underlying um hardware or investigating hardware to understand uh their neighbors really well and that was the way that for at least supercomputers um that's the way that you get a highly tuned system that can get get onto the top 500 for a small system and yeah maybe there is more blanket turnkey kind of solution. Yes. So David, you raised an interesting point uh in passing here, which is uh the the role of software.
We've managed to gloss over it up to now because I think largely of because of the work uh done by a number of uh people uh to develop the standards that we rely on now so that we can ignore uh the the complications so that we have you know uh infiniban verbs and MPI uh that's a very complex uh topic that that usually does uh and historically has uh depended you know, been in different states for different types of interconnects, let's just say it that way. And I think the community has done a great job of coming together and trying to hide that level of complexity from the average user.
It it's still there, but um I think we don't spend a lot of our time thinking about the details of of getting the hardware to communicate because of the work of the committees. C can you guys actually expand a little bit more on what you mean by price? Is it like literally the price of the hardware or is like Gary, you mentioned energy. You said it takes like a double the amount of energy. So when we're talking about how much things cost, what are what are we considering? Um you you know I was saying it was like uh a that it took twice as much compute on an Ethernet cluster as compared to an infiniband cluster to achie achieve the same amount to say achieve the same result on the linack test.
Um so um but if you're talking about price I I guess usually um when you're designing clusters there's uh there's kind of like the compute budget, there's a storage budget, and then there's an interconnect budget. And usually the interconnect on the more parallel systems, the the more parallelism that you're going to do on the system, the higher percentage of the budget should go to the interconnect. And so it can go anywhere from a interconnect with a blocking factor like a say 2 one 3 one even 4 one to what they what Allan called a full section uh uh bisection bandwidth where um essentially all the nodes on one half of the cluster could talk to all the nodes on the other half of the cluster without any kind of um blocking.
So that's where the cost comes in. Yeah, it's also probably worth noting that like when we're talking about networking and Infiniband and stuff like that, the the cables themselves even are more expensive than what you'd think of from a uh more traditional Ethernet network, right? for you know the cost is uh for the network and the computer resources is um especially for large systems which is where I spend my time uh doing research can be astronomical in a sense and then you have to worry about uh power costs and everything like that just to operate the machine. Yeah. Yeah. Correct. So, let's go back to a point that Gary raised that typically you'll make uh connections within a rack on copper and then maybe between racks or certainly longer distances with optical fiber.
Copper cable could cost a few hundred bucks. A optical cable can cost $1,200 depending on what it is. So um that when I say um a modestly sized cluster is easy is easier to build with a full fat tree full non-blocking folded cloth network if you want to be a computer scientist. Uh uh that that's one of the reasons it's the cables just cost you less uh and there are fewer switches involved. The bigger the systems, the more uh there's roughly a I mean basically identically a 2:1 ratio, but roughly when you fold into storage between the leaf switches and the core switches, the number of switches grows the cluster cost of switches is non-trivial.
It can be tens of thousands of dollars. And back to another point that Gary raised, the Ethernet uh sort of pre- slingshot Ethernet, classic Ethernet can be used to build clusters. And the reason that people were willing to put up with lower performance um you know uh more chattiness more latency and so forth in Ethernet was that the switches and cables are so much cheaper. Um so that's why especially a lot of the early Baywolf clusters were built on Ethernet. Um so one of the things too on cost um I remember when I first started in the group over at Sandia I was we were having uh we were really thinking about fault tolerance and power.
It was a really big thing. Resilience from fault u from a fault uh was important because we were seeing that systems were failing because the chip counts were getting up there uh to the point where we were scaling you know by the thousands of of flops uh each whatever cycle. Um uh I don't want to get into the whole when are we going to get access to scale thing but that that they were going through. It was always pushed out a couple of years. Um but uh it it did seem like the the um the fault rate of the actual devices was going to be so substantial that it wasn't the system would never make progress because it would always have to fall back on a checkpoint to get back up to speed and then go forward.
So when when you have more chips, when you have more devices that you'll have more failures and so that's something to keep in mind too, the maintenance cost and the uptime or downtime cost of of your system no matter what the size really. Yeah. As you add nodes, your meanantime to failure just uh shrinks drastically. And when you start looking at um positioning of your cluster, right, like um where at 6,000 feet here, um you go up to some place in the mountains like Los Alamos, it becomes actually significant compared to having something like closer to Berke Lake, which is closer to sea level. Yeah. Yeah.
That's interesting because um cosmic rays, believe it or not, it it really is a thing. Matter of fact, one of Matt's colleagues did his dissertation on just that. Uh I think on a study with a comparison between Los Alamos and might have been some Sandia stuff, but I can't recall. So, how does that impact? That's really curious. Is that is that just adding more errors or or what what does it how impact? It flips. It flips. Yep. Wow. Yeah. And there is of course error correction and uh and retries. Uh but overall NPI is still remarkably fragile. So um uh Jeffrey Fox did a lot of work on looking at the comparisons between for example map reduce and and classic MPI and uh looking for ways that the NPI uh standards could be made more robust and resilient to failure.
But pretty much if you lose a a node during a calculation, you lose the whole calculation no matter how many nodes are in it. uh still even even in spite of this kind of work I think we need more such work as we build exoscale systems with you know very large number of nodes the points that they they've raised has become very important and we still hear uh though they've done a lot to tamp it down we still hear a lot about the reliability of the largest of the exoscale deployments I will uh just hop in here for a second and say that this has been an ongoing uh topic in the NPI forum.
Um we're finally at a point where errors aren't necessarily fatal, which we got to maybe two years ago or something like that. Um but like even an out of memory error would just kill your MDI application. Um which seems a little insane in retrospect. But yeah, I should I should I should I think I believe uh Matt is on the MPI committee, right? Uh uh I am the Sandia representative for the MPI forum and in the open MPI developers community and I want to plug uh your your um your conference.
So, so um he's also a chair co-chair with another uh former colleague and uh friend uh for hot interconnects this year which is the le e standard I think it's the 29th year um and so I believe so um the uh so I'm the vice chair the chair is uh Taylor Groves out at uh Lawrence Berkeley National Labs. Yep. um working in Nurse and uh yeah uh if anyone has papers that they want to send over or we run a free virtual conference so it is uh there's a lot of good talks and there's no cost to attend cool nice thank you yeah so I'll just say on behalf of the community thank you because this is incredibly important work and it does require detailed understanding.
Yeah. So I have a question guys and I would like to hear from every single one of you if you don't mind and you know definitely it's going to uh you know change based on what the use case is. But say for example you're going to build a cluster let's say for a university like Allen's university like what would be like the perfect system like how big would it be what interconnects would you use? start with. I'll go first because I have no idea how to address that other than to say um my focus has always been on using the power that you have available and performance of a large cluster that is not utilizing its power cap is is is is kind of wasting money.
It's throwing money on the floor. So my research that I did was in uh focusing on uh power efficiency. So how well did an application or a system use its power envelope uh and how well could you reconfigure those uh into a lower power state or into a a smaller power uh window um or or a longer execution time to help load balance a system. So you could do that on any system. Um and you can make any system somewhat performant. I guess that's a subjective term. But I think that um if we start to look at systems instead of how much money can we spend and how much fabric can we have and how much you know of the new fancy stuff that we have out there, how complex can we make it.
If we start to look at our workloads and how can we tune them better and how can we really get that understanding of how well it an application or workload is utilizing the uh the compartment it's running in. That's that that would be my approach because I don't know how to answer it otherwise. So Rose, let me um uh back up a little from your question and and define the terms. Right. when you say uh university, let's just sort of describe what that environment is. Uh Gary's familiar with this uh as well. You guys do have a lot of customers in this area.
Um so university workloads are typically highly mixed in terms of uh they don't they don't have just one dominant kind of application or if they do it's uh you know it's uh maybe due to special funding circumstances of a group but uh uh you have a mix uh of workloads from very small core counts up to very large um and if you ask yourself I I should have prepared a a graph I have a graph where I show this the distribution of jobs versus core count is highly peaked at the low end of uh of small core count jobs. But if you ask yourself at any given time how many of my cores are occupied by small core count jobs versus large core count jobs after the schedule has done its work of balancing things out.
you find that it's sort of the 10 to 100 and 100 to a thousand uh bins that get the most populated. So yes, there are a lot of very small core count jobs, but each larger core count job takes more cores. So if you ask yourself how many of my cores are taken up, it turns out that the medium scale jobs uh and I think a lot of these environments will dominate. So then you have to ask yourself what can I do in designing my fabric storage CPU you know uh to uh get the most productivity out. Um that's where uh the questions I raised earlier come into play.
Um so I find you know with a we have about 750 nodes here and with a cluster on that scale um having you know uh non-blocking topologies just helps the scheduleuler do its job more efficiently and stuffing all the little jobs in to the corners where it can between the big jobs. Um that's not the only way to do it. Look at Bridges. It's a couple thousand nodes I think. Um they have a a uh overs subscribed fabric with non-blocking in I think it's not quite per rack uh maybe it's a couple racks expense they do nonblocking per rack and then overs subscribed between racks that means your scheduleuler has to fit those jobs those bigger jobs into a rack right so uh now uh I say rack but uh you know the densities are such that you really be talking about a switch pretty soon.
Um uh and let's say say you have a you know 40 core uh 40 uh port switch um with 20 ports going to worker nodes and 20 to uh to uh to upstream then or or maybe even overs subscribed then 20 nodes at uh you know 256 cores a node is a is a pretty big job. So these are the kinds of things that you have to factor in. And I think if you phrase it that way in terms of workload, then you can get out of uh the area of just talking about a university and look at other similar workloads. This might apply to a small bioinformatics firm or to you know a number of different settings.
When you get to the really large scale topologies, that's uh that's where the fancy fireworks come out. Right. Cool. So, it sounds like really what's driving this is really understanding your workloads, understanding what you're really trying to solve, not not coming out and saying, "How much can I buy, where can I, you know, how can I build this out?" Uh, it really is kind of reverse engineering it. Okay. Interesting. Yeah. And especially with the the the the rise of so much compute power in a GPU and the the need for data movement, um, I think that changes the whole equation. And I think Alan was mentioning that earlier.
It's a really important aspect. You have to understand your community and your workloads. What and and what's the future of uh of of the hardware that is coming out next because some of these supercomputers um what they do is they do a first initial install and then they plan out within a number of years, a couple of years or so an upgrade so that they don't just do that upfront cost every single time. they have a a more longer term path um so that they can get reuse out of them. Knowing what hardware is going to do for the next generation is an important aspect of understanding the system that you're building today and the longevity of that system for for for future.
I think the uh one thing that I can add here is that uh it also really depends on your user base. Um I was running a panel at XMPI last year and uh it was on the you know use of how are we gonna use MPI from GPUs. Um and we had a uh very impassioned uh member of the audience um saying you know why does this matter uh in in industry when we're using HBC we only program for the uh CPU because uh that's what's most cost effective. It takes too much uh developer time for a small company to port everything to the GPU and do all that other stuff.
And you know that there's a cost balance here of like how much time am I going to spend uh writing my code versus how much money am I going to spend on my cluster like on tour hours or whatever. Um and so that you know it for a university it's really you want to tailor it to the uh your expected work workloads and if you know they need something more than that there's resources like nurse and tac that they can get or that your users can get allocations on. Yeah that's a good point.
We've also been talking about this just in terms of on premises clusters, but the same considerations play out in the cloud and and the thing that distinguishes them is that of course the cloud tries to make its money by selling on demand and fundamentally that means um that they have to deploy at a scale that's sometimes much larger um but still be able to partition it up uh for individual workloads.
So um there's recently been a a paper pointed out by Glenn Lockwood which uh talks about the the infinibend structure of the RDMA structure of uh Azure storage for example and uh it's really a tour to force of uh of doing exactly uh the things you need to do to have a very large scale fabric uh deployment but still allows you to do things like uh take care of security considerations for keeping an enclave of that uh available only to that user. Um I'll see if I can find a link to that. So guys, how is it measured? Like how do you measure the performance of your HPC interconnect and actually I'd love to add on to that if we could because Allan you were talking about how there are HPC systems that are are connected differently, right?
And in rack we're doing at this speed. Uh so I'm always curious like how how would you know where or how would you figure out along with that? How do you monitor it? How would you know where there where your bottleneck is? That's kind of for everybody. So we have a lot of different ways of looking at this uh that I'm aware of at least. Uh you know there's lint level benchmarks. Uh I know nurse has collaborated on uh these large scale kind of noisy networks to identify um bottlenecks. What we focus on at Sandia is um proxy applications um which are um a small application that does something that looks like science um and communicates like science and uses the RAM like science but uh is easier to understand and uh doesn't have the test the like tour uh function or performance.
Yeah. Yeah. Matter of fact, yeah. So the Mantivo mini apps are are very much focused towards different aspects of computing different computing workloads. So certain workloads could be uh computebound uh memory bound um uh communication bound. They could be doing a lot of exchanges pointto-point, you know, all of the different sorts of um uh ways that uh you would behaviors. It's almost like they're different behaviors of scientific codes condensed down into a small kernel so that you can get these sorts of benchmarks and and we I some of the work when I was over at um Sandia I did uh some multigenerational studies and they have a beautiful set of test beds there a number of clusters of different some are heterogeneous and some of them are homogeneous.
this they're outfitted with some of the more interesting uh newer generation stuff that's coming out and we're able to see um at least to a small scale we'll do a scaling study still um to a small scale uh you know for 32 you know 96 nodes how an application or one of these mini apps the kernels uh perform as you start to scale and usually you can see how an application is going to scale you know just by a few different u you know doubling your size a few times. Yeah.
So with some caveats that many people uh don't handle the statistics of such benchmarking properly and that's another whole webinar we could do but uh the um the standard thing to do is to study your scaling and and practical terms let me point out if you're going to ask for allocation on a national scale supercomput you're going to need to demonstrate that your application is able to scale. Um, unfortunately I'm a little constrained on what we can say. We've actually been doing some work uh within my center um with one of the um member companies uh on improving the state-of-the-art for um benchmarking of uh NPI code.
Um u but um uh let me just go back to what I started with. It's easy to fool yourself with statistics when you're benchmarking. need to be careful, more careful than most people are when they're doing these scaling studies. Um, ultimately what the scientist cares about is how many units of X workload they can get through. Um but even simple tools like you know total view or debuggers can help you look at uh how you're using your memory how IO one of the things we have published from my center uh and we've talked about in previous webinars is that we have open source tools uh
at github.comnsfcac that you can use on pretty much any modern system that supports redfish to uh gather outofband statistics on things like memory power uh CPU power and uh we're working to fold in these bandwidth statistics so that you can do kind of a touchless uh uh overview benchmarking of your cluster you know how where is it spending time and effort um uh without having to code. Yeah, I think I think Redfish actually um was a utilized the power API which was a Sandia product that I was I worked on when I was there. And yeah, it's it is very important to get those out of band uh information and not interfere with the workload or else wait you're you're perturbing the system.
Um but it goes back to cost like you were saying earlier Rose were asking about cost. It isn't really dollar cost. It's science like like Alan had alluded to. It's the cost in science. How much science can I get done on this system? And how do I get it done very efficiently so I can get more science done? It's it's kind of as simple as that, I think. So, Gary, how often do you get researchers coming to you and saying, "Can you speed up my code?" versus just, "I can't get this to run." It does happen. It does. It does happen. Usually most of the issues are these days that I can recall like say in the last couple of months have been IO related.
Um and so it hasn't been with the interconnect more with the IO but it does they do they do. So I guess we're getting towards the end. One of the things that I was curious about is what because I talked to a bunch of customers um who are kind of looking for direction and I was curious what the what is the state-of-the-art right now if you were to recommend an interconnect and I know there's obviously depends on the workloads but what's what's kind of the state-of-the-art what's and Allan I know you you may not be able to say and dive into what you're talking about but what's if we were to go to offtheshelf what would that be listen I think it's more competitive than people realize uh like Gary said most people's starting point is infinib band.
But let's face it, Infiniband is expensive and touchy. Um, I've got both an Infiniband and Omnipath deployment, and the on path never gives me any trouble. Um, the Infiniband I have to be careful with every version update, and there seem to be lots of ways that small communication errors can interrupt performance. Um the I'd mentioned the switchless fabrics uh you know the rockboard style ones you just put a different PCIe card you don't have either infiniband or omnipath but it is limited in the number of nodes nodes you can connect uh the great advantage of that is it's very little latency I mean essentially no latency um well not none but instead of pci backplane latency at that point um then uh we we also have the PCI interconnects, the the uh the uh composable computing fabrics.
So, you can buy a box of GPUs and shove them in. And depending on your workload, again, if you need those GPUs to communicate or not communicate, you could save a lot of money if you don't want to spend a lot of money on interconnects by just having a box of GPUs running them separately. Uh so, uh you know, all of the above. Okay. But I think most people do start, as Gary said, looking at infinite band fabrics, and that's a good place to to get your feet wet. I think uh what was it? Uh with Melanox now being uh owned by Nvidia. Um and then uh Omnipath is spun out of Intel.
There's kind of a hardware dependency that you that is a little bit baked in to some of these. I would say affinity. Yeah. Yeah. I I I worried the same sort of thing when XYlinks and Altera left the marketplace and got consumed by AMD and uh Intel. Same sort of thing. They're smart. They they know that's going to be the the optimization path. And they also know that the um the ecosystem and the people that are out there that can do those things are very limited. Uh even though Xylinks was trying to create the tools for an ecosystem of design engineers to make their own uh fabrics.
Um so in a way um you're steered you're you're you're always a slave to the uh uh the industry and the movement in the industry. So, your decisions of what you get could have to do with who owns what, which is a sad reality, I guess. Yeah. Running out of things that Nvidia hasn't bought, right? Yeah. Right. Well, going down that path, do we see GPUs as kind of being like where this is all heading? I mean or or is there still a need for CPUbased compute uh just for pricing or or cost? I think it depends on the the workload and and yes that we definitely do see a need for that with with the amount of uh small compute uh and large data movement um that you see in some of the machine learning uh codes and training.
Well, and it also depends on your users, right? I alluded to this earlier. we have, you know, a bunch of people working on GPU codes, but there's a lot of people that don't want to deal with that and don't want to deal with the added complexity to their science of having both the CPU code to manage everything and then moving stuff over to the GPU and moving it back, doing your communication. Um, we're trying to figure out how to do MPI from the GPU, but that's very in its early stages and has not been standardized and we're working on that in the MPI forum, but it's not a reality yet.
Yeah. Yeah. I think it's a little informed by the topologies of the the big clusters too, right? We've seen um let's see, Frontier is uh four GPUs per CPU if I recall correctly. Um then uh I think the new uh AMD CPU plus GPU systems are just being deployed where you have essentially uh you know all on a socket uh or all within a you know a small portion of the back plane. So my controversial take that can never I can never get hardware vendors to agree to is that when they build uh national scale what really big system they should sell little units of that system to a bunch of universities just identical because then we can get the most out of the software development.
Yeah. But I can never get any of the hardware of interest to agree to that because they lose their shirts on these big systems. Yeah. Well, we are coming up on the top of the hour. Any final thoughts from anybody? Well, I I want to thank Matt for doing this at the last minute. I just called him an hour before this to uh to uh see if he can make it. I'm glad he could. Uh Matt's a good friend of mine here in Albuquerque and we uh Yeah, it was good to see him. Absolutely. Thanks for having me. Sure. Yeah. Love to have you again. This is great.
Yeah. Thanks, Matt. Any other final thoughts? Everybody else good? We talked it out. Well, you guys make sure that you like and subscribe. We are here every single week, same time, same place, same channel. We will be here. If you guys also have um suggestions on topics for webinars, things that you want to dive into, um please leave those in the comments. We want to read them and we want to serve our community the best that we can. So, thank you everyone on the panel for being here. Thank you for watching. You are very amazing. And we'll see you um next week. Same time, same place.
See you. All right. Cheers. [Music]
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.