
Research Computing Roundtable - Turnkey HPC: What is HPC?
Webinar Synopsis:
Speakers:
-
Zane Hamilton, Director of Sales Engineering at CIQ
-
Krishna Muriki, Senior Software Engineer, KLA
-
Patrick Roberts, Technical Director, Skyworks
-
John Hanks, System Administrator, Chan Zuckerberg Biohub
-
Gary Jung, HPC General Manager, LBNL & UC Berkeley
-
Glen Otero, Director of Scientific Computing, CIQ
-
Forrest Burt, HPC Systems Engineer, CIQ
Note: This transcript was created using speech recognition software. While it has been reviewed by human transcribers, it may contain errors.
Full Webinar Transcript:
Introduction [00:00]
Zane Hamilton:
Good morning, good afternoon, and good evening wherever you are. We welcome you back to another CIQ round table. We had one of these a week before last, and it went over really well so we invited a group back to continue that conversation and see what else we can get from people who are in the industry and actually have real world use cases and talking points. If we could go ahead and bring in the panel.
Welcome back everyone. We have added some new participants this week too. I think last week, everybody got to meet Patrick, John, Gary and Greg as always. So I think we will start with the new guy in the room. If you would go ahead and introduce yourself, Krishna, and let us know who you are and where you are from.
Krishna Muriki:
Hello. I am Krishna Muriki. Right now I am employed with KLA Corporation. But I go a long way with Greg Kurtzer and Gary Jung, both of them here on the panel. Before this, I was employed at Lawrence Berkeley Lab. Greg hired me into that position at Berkeley, and Gary was my math manager and guide. I know the panel very closely. I have worked with Patrick on Slack. I learned a lot from Patrick. Glen, a good friend of mine. I am glad to be here.
I am about 10 months old at KLA Corporation. In this organization, we make wafer inspection devices. In the semiconductor industry, it is very crazy right now with all the chip shortages. The demand for our product is really high. It is a little bit crazy. We are having to stretch ourselves in terms of our HPC designs and HPC technologies. We are adopting to meet these growing crazy demands. This topic is really interesting. I am glad to be here.
Zane:
Very nice and welcome. I am going to go around the screen real quick, just introduce yourselves very quickly again. I think because you introduced yourself last night you do not have to go into a lot of detail, but I will start with you, Patrick, because you are at the top of my screen here.
Patrick:
Howdy all, my name is Patrick Roberts. I currently work with Skyworks. At Skyworks, we design all sorts of different types of chips for all sorts of different things. So sorry, that is rather vague, but we are in all sorts of RF chips, analog chips, digital chips, big and small. I have been around for quite a while. They used to be Rockwell long, long ago, so some of you may be familiar with that. It depends on your age I suppose. But I have been in HPC specifically for a little over 20 years. I have been in the IT industry for 29 or so years. That is about it.
Zane:
Next we have John, aka griznog.
John:
Hi. John Hanks, I have been doing this for a long time. I do not remember what I said last time when I got introduced so I hope I do not contradict myself this time. I was thinking about the topic today and I think maybe I am inherently unqualified to answer this question or even try to answer this question because in pondering it, I think all I do is manage workstations now. They just happen to look a lot like a cluster because of the way I manage them. I do not know that I even do any HPC anymore.
Zane:
No, I think that is great and that is something that is going to be interesting to see. What is HPC to everyone else? I will move next, kind of skip over Krishna because we had you. Greg, you are in the middle.
Greg:
Hi, everybody. I feel like, who is in the Brady Bunch in the middle? Was that Alice?
Zane:
Jan. No, I am kidding.
Greg:
Jan. Thanks. But hi everybody. I have been doing HPC pretty much my whole career. Started with a degree in biochemistry and then failed to continue getting my PhD because I got more interested in computing and Linux and open source than doing science. I have been really focused on HPC and scientific computing ever since. As Krishna mentioned, I have worked with Krishna and Gary for ages at Berkeley Lab. It is great to have this middle line across the board as the Berkeley crew.
Zane:
We will just call it the Berkeley line there. The next person from Berkeley, Gary.
Gary:
Hi. My name is Gary Jung. I manage the institutional HPC computing at Berkeley Lab. Greg and I started the program back in 2003 and it has been going great since. I also have a dual appointment with UC Berkeley. Also, I manage the HPC effort for the institution down at UC Berkeley. Krishna, Greg and I go way back doing this.
Zane:
Glen, you are up next.
Glen:
I have been at CIQ for 16 days. I am the VP of fricking sweet ideas at CIQ. My background was a doctorate in immunology and microbiology. I have a postdoc in HIV research. Then I left academia and fell into HPC with my second postdoc at the San Diego Supercomputer Center. Since about ‘99, bioinformatics and HPC have been my livelihood. I have known Greg since about 2000. So I blame him.
Forrest Burt:
Hey, everyone. My name is Forrest Burt. I am a high performance computing systems engineer at CIQ. I have been with CIQ for about a year. Before that, I was the student HPC system administrator at Boise State University while I was getting my bachelor's in computer science. I got my start in the cloud and in computing and stuff like that in the early 2010s, messing around with the tech that I had available and cloud gaming servers. I am really pleased these days to work with a little bit bigger systems, but I love HPC. I got super interested in it while I was in college. I am really pleased to have been working at CIQ and to be working on everything that we do with Fuzzball and that type of stuff. So great to be here.
Patrick Defines HPC [00:07:56]
Zane:
Thank you, Forrest. Jumping back into the question of what John has started to allude to a little bit, I think that the definition of HPC is going to be a little bit different to everyone here. I would like to go around and ask: what does HPC really mean to each of you, and just tell us what that means. Tell us a little bit about maybe where you got to that point and how you got there. I will start off with Patrick.
Patrick:
I would say, what does HPC mean to me? So HPC is any service which is designed specifically to perform a set of complex tasks, which I know is a very broad definition, but it is very true. Because I have dealt with HPC in all of its various forms from Cray to working on loose grid systems to working on top of very close coupled, very tight systems with InfiniBand and MirrorNet, and just all sorts of different interconnects of all kinds, and it really just depends on what tasks you are trying to accomplish, and what you need in order to meet that task out.
There is a blending of HPC for high performance computing and then there is now high throughput computing, and now there are all these other subvariants, but in my mind, it is all HPC; it’s just with different tasks in mind to accomplish, and different hardware and software solutions to accomplish those tasks. So for me, it is just one really big bucket of designing clusters and computers to meet tasks' needs.
Zane:
No, that is great. Thank you, Patrick. And John, you started down a road of talking about desktops and that is what you deal with today. I know it may have been a little bit tongue-in-cheek, but in reality, that can be something that is really the case. Tell me about your definition of HPC.
John Defines HPC [10:09]
John:
For me, HPC, the primary thing there is high performance. And for it to be HPC, it needs to be high performance, which is why a lot of what I do now is not HPC because I have people getting this massive pile of GPU hardware, memory and cores, then sitting there typing in a Jupyter notebook on that system. It is as far away from high performance as you could ever hope to get, somebody checking out a $7,000 GPU to type and hit enter at a command light prompt.
I think HPC, high throughput computing, those are really the two main designations. And then for me, it is getting even further away from high throughput computing, where we run some embarrassingly parallel workloads, but the most important thing in our environment is on-demand for people interacting. So we are pushing that long tail even further out by pulling in people that do not really need a cluster. They just need a well-managed desktop workstation that has a lot of memory.
Zane:
Very nice. Thank you very much. Krishna, how do you define HPC?
Krishna Defines HPC [11:28]
Krishna:
I think it is changing. I think the definition of HPC is changing. In 2005, when I started my first job in the US, it was at the San Diego Supercomputer Center. The first machine I worked on was from IBM; it is called DataStar. It has a very specialized interconnect, very specialized work servers all put together, but that was not the high-end machine back then. The San Diego Supercomputer Center had a good size machine, DataStar, which is how the HPC machine used to be. The applications that used to run on that application were very optimized, tuned for the IBM architecture, which we used.
These days, I think it is not the case anymore. A lot of technologies have commoditized and interconnect InfiniBand, as it is treated as a commodity. A lot of servers that we use, even the top machines, the top 500 machines are commodity servers from Supermicro and things like that. I think it is morphing. The definition for HPC is becoming a lot more commodities and not specialized machines anymore.
The employer who I work for right now builds stream computers. All that we have is a lot of data streaming at us. We have really high-definition cameras looking at this nanoscale layer width, transistor width, and things like that. It is really high volume data, just a fire hose of data being shot at us.
The biggest machine that my employer built is just three racks, not hundreds of racks or even tens of racks. It is just three racks, which is the biggest that we have. It is all commodity hardware. We do not have a specialized server design team at KLA. We do not have a specialized architecture team at KLA. We all take white boxes, Supermicro, InfiniBand, Commodity, and just try to build. I would consider what we are doing at KLA is HPC because the data volumes that we are dealing with, if that is the definition, does come under HPC. But hardware wise, it is all off the shelf. It is changing. HPC definition is surely changing, is what I would say.
Zane:
That is very interesting. Talking about commodity hardware versus very specialized hardware that it used to be. I think, like you said, things have changed.
Patrick:
I do not think that is inherently a bad thing, though. I really think that commoditization of the hardware has brought HPC into the hands of so many more people. It is actually a very good thing that you can build a super equivalent with off-the-shelf parts. I did it for a very, very small but, one could argue, quite important X86 design company, where they did not have the money to go out and really invest in a massive super to solve their complex problems. But instead, I was able to put together desktop machines, literally on a folded piece of metal.
And I would go to Fry's, rest in peace, and purchase off-the-shelf hardware and literally take it back and have one of my guys put it on a folded piece of aluminum and bolt the bits to the folded piece of aluminum. I would have bread racks from Sam's Club with 44 machines per rack in an open air data center, because it was designed originally for mainframes. It was designed to have an open air concept. And roll in the bread racks of 44 machines. We could do it and add thousands of cores at a fraction of the price that it would take for us to go to IBM, Supermicro, Dell or anybody to have it done.
So just the ability to do that, it allowed that small, little 100-person company to compete with the likes of AMD and Intel. Not necessarily effectively compete, as they got bought by Intel last year, but they were able to compete. Just that alone, that commoditization of hardware, is not a bad thing in my opinion. Yeah, it is really cool to work on those really highly specialized systems, but it also is amazing what bringing that capacity for computing to the general populace can really do.
Zane:
Before I go to Greg for his definition of it, coming from an enterprise background, I have watched people build commodity infrastructure to solve enterprise problems that probably should have been HPC in the past. It is interesting to see the merging of the two worlds where you see HPC becoming more commodity hardware. Maybe those workloads are moving from just general enterprise into an HPC environment where they actually get handled in a much more efficient fashion. I have been excited to see that merging. Now I go to you, Greg. What is your definition of HPC?
Greg Defines HPC [17:08]
Greg:
Before I even mention that, I just want to respond to Patrick's point, I completely agree. We have seen this happen so many times in research organizations, especially governments and academia, where in many cases, people that have never done system administration before are even buying clusters. They are not managing them very effectively, but they are getting research done. Because of small budgets, they are actually able to do very good research on it.
At some point, typically they will graduate and be able to have proper support for those systems. It is not a grad student or even an undergrad or postdoc who lost a bet and now has to go maintain that system in many cases. They can actually reach out to somebody like Gary's group at Berkeley and LBL, to actually solve that problem and manage it appropriately. But hitting commodity has lowered that barrier of entry so high that it’s been fantastic.
Now, to get to the question you asked me, Zane, I have been doing this for long enough to have a very strong opinion of what high performance computing is, or rather was. It was tightly coupled parallel processing at massive scale. I remember times where unless you can prove that your application could run effectively at massive scale, you did not get time. You were actually down revved in terms of how much time you can get and what priority you get to run in, unless you can demonstrate that level of efficiency.
Now, we called it the long tail for a long time, which was all of the scientific applications and workloads that are running on these systems, which are not that massively traditional high performance computing profile. What we have seen over time, at least what I have seen over time, is that long tail has just continued to grow. Griznog, you kind of alluded to that as well, that the long tail has now grown so far out that the long tail is now becoming the dominant form of workload on many of these systems.
I have had many, many arguments over drinks at supercomputing on what is high performance computing, and are these sorts of applications actually high performance computing. It took me a long time to actually get to this point in my maturity as a person to say that high performance computing is now more than just those massive tightly coupled applications. It is anything that is going to spin or peg some aspect of that hardware at 100% and then literally bottleneck based on that subsystem of your hardware.
Whether that is CPU, whether that is memory IO, whether that is bandwidth, you are going to peg something on your system. That is where I am in terms of the definition of high performance computing. It could be a single thread hitting one core, or even a hyper thread on your CPU, but it is pegging something, dang it, somewhere, which is where I draw the line.
Zane:
I like it. Efficient utilization of resources, which is all I am hearing.
Greg:
I am not sure if it is efficient in many cases, but it is definitely useful.
Zane:
Pegging it, resource, whatever. Gary, your definition. You have a fairly large environment to deal with.
Gary Defines HPC [20:53]
Gary:
It is interesting when Greg talks about that, it brings back a lot of memories. It was exactly the way he said, where we are looking at really big usage of the system. We do take a simplistic view of what HPC is. We work with a lot of researchers, large team science, and then also individual researchers. A lot of individual researchers may have a very powerful workstation, they may have a laptop. Then at some point, whether they are doing data acquisition or simulation, they are going to outgrow it. They have to move to something.
What is that something? And then for us that is where we define the HPC. Instead of coming up with one offer for all these people, then that is our opportunity to move them into an HPC infrastructure. And then I think everybody's kind of covered that, but there's enough convergence in the field so that now when you describe an HPC infrastructure, everybody knows exactly what you are talking about.
Greg:
One thing I just want to follow on what Gary said, speaking of memories, we did see a lot of people that started off on a laptop then went to a powerful workstation. Then we initially, when we first started the HPC project at Berkeley, it was around this concept of something called mid-range computing. Mid-range computing we defined it as everything between a workstation and a giant facility like NERSC, and everything in between.
NERSC at that point was limiting people unless they can prove a reasonable amount of performance and efficiency in terms of running on their system. They did not want applications running on that system unless they were efficient on that system, because it was such a big, expensive resource.
One of the big motivations we had for this group was this idea of mid-range, and how do we solve these mid-range problems? To me, it has now turned into what I call sweet spot computing. This is where most people are focused on. For example, in high performance computing, this is the vendor's bread and butter. This is where they are selling the most systems. It is why we are talking about turnkey HPC, honestly, because it comes right back to that. It is the problem set we actually really need a solution for.
Gary:
Yeah, that was a huge gap. I think essentially when we did that early, we helped to define what is known as the mid-range computing segment within the DOE.
Greg:
I think that actually even articulates how HPC was changing even then, Gary, without us even really coining it or thinking about the nomenclature. We were changing what HPC was even then. It started off being those giant systems and then it was moving towards mid-range, and it has been still following in that trend and in that direction ever since.
Zane:
That is very interesting. Thank you very much, Gary. Next we will move to Glen. You have an interesting perspective on this. What is your definition of HPC?
Glen Defines HPC [24:20]
Glen:
When I started in this field, HPC was tightly coupled, pretty large scale machines, Right? You are in that elite interconnects. But when the life sciences started adopting the Beowulf architecture, it rarely had those little latency interconnects because it did not really need it. They were coining at high throughput computing on ethernet.
It is not really HPC. It is just high throughput. But in my head, I was like, same architecture, it does not have your low latency interconnect. It is a different type of parallelism. I still thought that it had to be something like HPC.
Then in my head it evolved to, so, how do I lump those two together? Then that became research computing. Research computing was all these folks, whether it was molecular dynamics or just DNA alignment, and embarrassingly parallel or tightly coupled, it was all research computing. Then the enterprise was doing it. Now what do I call it?
Then to John's point, I have worked on machines with all these GPUs that are running a Jupyter notebook, but you better have that RAPIDS library and that Python tuned, optimized really well so they get the most out of it. At my last employer when we were working on my title, which was not VP of fricking sweet ideas, unfortunately, we talked about, well, why not scientific computing? You are trying to get science done, really, whether it is Johnson & Johnson trying to figure out how to design their next shampoo bottle or if it is COVID genome analysis I have come to call everything that we have just talked about scientific computing. If you had to nail me down to HPC, I would be thinking of systems that try to look like the top 500, tightly coupled, the scale of the top 500. Anything smaller than that is still HPC. I am not sure what my mid-range HPC red line is, but that is how I view the world these days.
Zane:
Thank you, Glen. So, between what Greg had talked about and what you talked about, I used to be on a mid-range team. And mid-range to us just meant the Solaris team and one of our guys called himself the VP of Blinky Lights because we had the data centers.
Glen:
Wow, that is a good one.
Zane:
Not quite as cool as sweet ideas, but pretty cool. Forrest, what is your definition of HPC?
Forrest Defines HPC [27:16]
Forrest Burt:
I would define HPC surrounding the efficiency, which we are trying to achieve with HPC. We have hypothetically a single commodity server that we are trying to do some traditional HPC workload on, whether that be simulation or just to roll them both into one thing, some type of embarrassingly parallel high throughput type of job. If you only have one server, you are not going to be able to get the performance you would need out of it in order to actually efficiently iterate on the work you are trying to do.
To me, HPC comes down to being able to do these really, really complex things in an efficient way, regardless of what the system underlying it ends up being, whether that is a bunch of EC2 instances, which is a traditional tightly coupled cluster, whether that is a cluster that is more meant to provide as many GPUs as possible, just to do mass model training, that type of thing. I think ultimately it comes down less to the comprising of the system itself. Are you building the system to gain performance beyond what you would get with the base component of this system? In that case, I would say you are doing HPC.
I would also touch on the long tail concept because where I very, very first entered this was more so in those tightly coupled jobs. Especially at my last institution, I started to see a lot of researchers coming out of what I felt were non-traditional places for HPC. We were not just having genomics people or simulations people, that type of stuff. We also had education or people from the education department that were trying to do mass data science crunching for their research.
We had people on the high throughput side. We had people that were taking LiDAR data and doing huge amounts of analysis in a high throughput way. I definitely think that in addition to being about efficiency, we are definitely seeing this kind of enterprise. Some of these other use cases realize that they are trying to do HPC without realizing that they are doing it. We see maybe some of the muddying of the waters here as some of these nontraditional entities enter the HPC sphere.
Have HPC workloads changed over the years? [29:35]
Zane:
Absolutely. That leads into one of my next questions, Forrest. And thank you for that. I think we all alluded to it at some point, but have the HPC workloads changed over the years? We talk about the infrastructure getting better, and the machines getting better. Obviously, they are getting more and more dense and being able to be more capable. Have the workloads changed over the years? And I will throw that out to anybody who wants to take it.
Patrick:
100%. Absolutely the workloads have changed. You have some of the traditional workloads, which are still around and NPI-based ones. We touched on this. I can not remember if it was in one of our dry runs or in the actual meeting last week, but we definitely touched base on who is still using MPI. What are your percentages of your workloads of different types and that sort of thing. It was definitely eye opening.
I expected it in some ways, but it still was very informative that the mixed workload is really here to stay. Even in very specific industries, the amount of workloads your singular cluster or clusters are supposed to be able to support is massive. You are supposed to be able to support GPU-based AI training or ML training versus being able to run ultra-fast single-threaded performance or highly parallelized jobs. You are supposed to be able to run it all in one.
All of those require different pain points to overcome in building the HPC clusters. I think we are expected to, as designers and administrators of those environments, we are supposed to be able to provide for all of those in a one-size-fits-all. It is difficult to accomplish.
John:
I think to some extent, the commoditization of the compute part of this has made HPC simple to do on the compute side. Where the actual interesting stuff being done now, and probably more so in the future, is on the storage side. Storage is actually becoming the critical thing about everything we do, regardless if it is HPC high throughput. At least for life sciences, a lot of on-prem clusters and on-prem infrastructure exist today because you just simply cannot move the data sets around. The compute has to be brought to the storage. The compute is more or less secondary to just getting all your data in one place where you can work on it.
Krishna:
Not only that, but the way the compute is being orchestrated or designed is quite different in the enterprise. Enterprise is dealing. I know of use cases where they are dealing with data pipelines at the rate of multiple tens of gigabytes per second constantly coming. That is the data volume they are dealing with. The workflows and applications they are running, they do not have any NPI. No NPI at all. Nobody in that organization knows anything about NPI. So yeah, NPI is not the only HPC. The data volume'-s are a lot higher than many organizations deal with, but these organizations do not have anything NPI.
John:
Lack of NPI in the enterprise pushes me into an old man rant moment about a negative connotation of HPC. In the early days of HPC, I remember about 20 or 30 years ago when I was getting started in this, there was an elitism aspect to HPC. You had to be able to afford to do it. It was a very elite club and not everybody got to join that club.
That attitude carrying forward over the following 20 years from when supercomputing got started, led to us having to deal with things like Hadoop, Spark and Kubernetes and all this just nonsense, complicated, extra complexity, poor performance because that is the stuff that came out of the enterprise. At any point during that, we could have just said, “Hey, enterprise people, how about you help us write a better NPI and let's do this right, efficiently and simply, with a simple scheduler.
Patrick:
I do not think it was that it was not offered. I think that it was not necessarily accepted amongst industry, because so many different corporations were trying to capitalize on that market. Or, well here, let me enter my proprietary standard into this ring. Let's go ahead. And they did not want to do an open source either, generally. They wanted to really make it very tight and hey, well, now that you have entered into our ecosystem, you are stuck.
That is what came about, I think. Then of course you had the competition, as you said, where now we have got Kube and now we have got all of these other competing, not exactly optimal frameworks to put it lightly, to do a lot of the same workload that could have been much better implemented in HPC.
Zane:
Let me ask this, Patrick and John, from an enterprise perspective, if I am looking at this, the way most of those other languages and other tools came out was to be simple and fast. If I look at NPI, it is probably a little more complicated to program correctly and to get it as efficient as it should be. Was it something that from an enterprise perspective, people did not want to take the time to learn how to code correctly for NPI? It was just easier to use some other framework?
Patrick:
It was a major barrier to entry. Because of that barrier to entry, whether it be I did not have the hardware to have that high speed memory interconnect between the nodes, or I am an undergrad, I cannot afford to get time on a super at the lab, or whatever it was. There is always a barrier to entry. Or if it is Intel NPI, do I have to buy into the entire Intel ecosystem? Do I have to go buy an Intel compiler back in the day? Does that mean that now that I am running on my AMD chips, is it going to be slower because Intel spiked the compiler? Whatever it is, it depends upon the environment, but there was always a barrier to entry. Then of course, you had open NPIs; now we have competing standards.
Greg:
I am going to push back a little bit because I think It is very use case specific. The use cases and the applications we were solving in high performance computing were very different from what enterprise was really even looking at solving. Now, the MapReduce kind of focus there, I do agree. We could have done better in terms of cross-pollinating between Hadoop and then Spark and then HPC NPI, NPIO, and other things.
We are starting to see that also happening in AI, starting to. This is a few years old now at this point. But things like, how do we do training? How do we optimize training? How do we start parallelizing training? If anybody remembers some of the early training frameworks, it was not running NPI. They created something from scratch, almost like NPI did not exist. It was a duplication of effort.
To the point, Patrick, I think you made this, it may have been Griznog, but if we would have come together right at the beginning, in a better way, more open way, I think showed more interest in collaboration, I think we would have seen very different outcomes. This goes back to, I joke about this, in all of the organizations I have seen and worked with that have HPC-focused people, traditional HPC-focused people, there is IT and then there's HPC.
Even if HPC is under IT, they do not talk; they do not work together. They do not leverage each other because the technology is both so different and they are solving such different problems. There is also just this legacy mindset on both sides. It is not until fairly recently, in my mind. Containers were the Pandora's box that really started blurring this. It was not really until containers did anyone in HPC even care what the heck was happening in enterprise?
We took off our gloves at that point. On the enterprise side, it was not until just recently did they start looking at HPC saying, how do we do training faster? How do we do better analytics? How do we get more insights from our data? That to me was the crossover point. Now we need to actually all come together and realize we are solving very similar problems. We need to stop innovating on our own and start actually working together and solving problems together. We need to be a big, happy family again.
Zane:
Need to be, yes. Gary, I think you were going to add something to that.
John:
What do you mean “again”?
Forrest Burt:
I just wanted to riff off that, a little bit of a side jump, but back to the topic of how HPC is changing. Something that I really see is the hardware side of HPC changing. It seems for a long time, we had InfiniBand, CPUs, GPUs, and that was kind of the basic components of a cluster as far as the compute part. We start to see now that there are a lot of companies that are starting to invest in creating new solutions around some of these very specific use cases.
Now we see for AI, we are only 10 years off of Alex Krizhevsky, I think I pronounce his name right there, training AlexNet on two GTX 580s. And now, 10 years later, we already have companies that are trying to replace the GPU for AI training because there are better ways that can be done than just on GPU. Some way that I really see it changing is on the hardware level.
We have got new ways of training AI, new much faster interconnects than we have ever had available before, new storage technologies. It seems to me like we are really in a renaissance of HPC hardware, the use of PCIe as a broader interconnect than just to put accelerators into a cluster and stuff like that. It seems like a Renaissance as far as HPC hardware goes. Our clusters are really starting to change from just the basic CPU interconnect GPU type of thing with FPGAs specialized chips.
Zane:
Forrest, I am glad you brought that up because that reminded me of a question I had for Gary last time, when he made the statement that his GPU usage has gone up by 3X over the last few years. What is driving it? Is it something that is more hardware specific since it is difficult to just go purchase whatever you need because it changes all the time? Is that what you are seeing? Is that what is driving it?
Gary:
For a lot of people, they are working with workflows. Somebody wants to put together or string together an infrastructure with a web server and a database and do some computing and move some data, and it is pretty easy. The term I like to use is frictionless computing, because they are able to do it themselves without having to, say, wait for a network drop in a data center and then for somebody to set up the machine and all that. It is a lot faster.
Then probably the other thing is it’s just really easy to prototype things. New tools like AutoML and SageMaker make it really easy for people to get started using machine learning. It is interesting. We will onboard people onto the cloud. We will kind of prototype it using AutoML or something like that and they will get going. Then sometimes what happens is they will find out it costs a lot of money and then we will go back. Now they see how efficient it is and how this helps their science, we will go back and transition onto our on-prem HPC using maybe slightly different tools.
Vector Processing [43:16]
Zane:
Interesting. I was just curious because it is interesting to see how it is taking place. I talk to a lot of different customers who are HPC, who are trying to, like you said, try things out on a cloud before they actually make the investment to see if it’s going to work for their needs or do what they expect it to do. Thank you.
We have a question from Mr. Ignite, so before we pop it up, I can go ahead and start reading it. Nobody has really talked about the vector processing as a part of HPC buildup. The example that is given is a medium scale IOT collection and analysis of billions of devices generating tens of bytes per second, driving anomaly detection, pattern recognition, and other applications. It is an interesting topic. I do not know if you guys have any thoughts on that.
Real-Time Stream Processing As Part of Cluster Work [44:01]
John:
I would be curious to hear if anybody has ever done anything on a cluster that was real time. The closest I have ever got was a group that wanted to produce a weather forecast every day. They had data arrive and then they had a deadline to get stuff out, but I do not recall ever having anybody who wanted to do real-time stream processing as part of a cluster workload.
Gary:
Actually, Krishna, Greg, and I could probably have an example of where we are currently building a system for a gamma ray detector and the system is going to be the first level trigger of the data coming off the detector. It is going to be a Kubernetes cluster. We are going to run different containers depending on what is going to be analyzed. We are using, instead of FPGA hardware, they figured It is going to be a lot more flexible to do with a cluster. It is currently being put together, engineered and put together right now.
Detector Research [45:12]
Greg:
I figured Gary and Krishna would have some examples here. There were multiple different examples of people doing detector research. Basically, you can think of a detector as a very simple idea of one. You have some sort of reaction happening within a spherical enclosure. Then you have a whole bunch of detectors all the way around this spherical enclosure. It is monitoring all of the different types of rays and what is happening at the center of this sphere.
All of these different detectors are creating streams of data. Then these streams of data have to be computed in parallel in real time, so you can see and visualize what is happening within that experiment. Sometimes these are actually in real time, but in many cases it is near real time. Which basically means now you actually have to look at it and then decide, do you want to rerun the experiment? Do you have good data? Then feed that back into your process flow. It needs to be real time or near real time.
As Gary just mentioned, if you are running it through Kubernetes, it is an extraordinarily different architecture than what we have been doing in high performance computing. Without even knowing the specifics on it, I would say probably it is a bunch of services, which are running and doing the equivalent of an inference on each one of those data streams coming in, and then computing on those via GPU. This is like a performance critical service that is running.
Kubernetes is a good architecture for managing services. I am curious about the efficiency that you are getting, because I have actually had really bad results with regards to Kubernetes. Just to go on a quick tangent real quick, we actually were talking to some people that were doing something along these lines with Kubernetes. Because Kubernetes is so big and bloated, it was utilizing about 15 to 30% of the underlying hardware resources just to run Kubernetes.
Now let's scale this up. Imagine buying a 5,000 node cluster or even a 500 node cluster and 100 nodes out of your 500 nodes is just running your scheduling system, your orchestration system. You are basically losing 20% of your system to run just for Kubernetes. At a small scale, I think Kubernetes would do it. As soon as we start scaling bigger, we have got to come up with better ways of doing that.
DDN Appliance For Real-Time Stream Computers [47:51]
Krishna:
Another approach I have seen in the industry for these real-time stream computers- is a traditional way If you look at a DDN appliance, a storage appliance, they usually come in pairs. They come with some fault-tolerant resilient designs using basic technologies like Corosync, such that services can flip from one controller to the other controller automatically if there is a failure in top controller. Luster services can move on to the bottom controller.
That is the basic components that anyone would use for fault tolerance. I am sure that even Kubernetes is built on some base components like that. In the industry, I have noticed, if you build a real-time streaming computer, where your machine goes into your production pipeline and they need real-time feedback of the quality of the product being manufactured, and if this computer is down and that quality is not being generated depending upon the value of the product you are making, it can cost multiple hundred thousands of dollars per day of production loss. That is very costly.
This real-time computer giving the feedback on the quality of the product is a very critical component. When you are designing such a computer, fault tolerant designs, having backup components is so critical and things falling back automatically from one server to other server is such a critical thing in your design. If you do not have it, you are costing hundreds of losses to your company, to your product. That is something that I am noticing heavily.
Not everybody in the enterprises is adopting Kubernetes. They are just trying to accomplish that using the proxy Corosync services, which have existed for a long time and trying to put together everything, like having redundant switches, redundant network cables, run from servers to switches and have services for failover from one switch to the other switch. If there is a switch failure or a network cable failure, those are another tangent and extension to HPC design that I am seeing in the enterprise, which is very critical and needed when you are building a real-time computer for these use cases.
Zane:
From my past experience, most of what we were trying to do was as much analyzing of data at the edge as we could so that we were not sending massive amounts of data back to be processed or analyzed. We were looking for that anomaly and then whatever data was driving that is what we were sending back. From a real-time usage, I feel like just because we were trying not to send a lot of data that may have been irrelevant, we were doing a lot of pre-processing at the edge. Then only moving back what was necessary for the next type of action or the next time event that needed to be taking place.
But that is an interesting question when you are talking about massive data. Like in Greg's example, it is not really an edge use case, because everything is right there together. You are trying to do it in the place that it is taking place. From an IT perspective, when you have billions of sensors out, that amount of data streaming back could be something that would be overwhelming. A lot of that at the end of the day just gets dropped on the floor and you never even look at it. There are a lot of different ways to approach it when you look at IT. It kind of depends on what you are trying to accomplish, in my opinion.
Customized Hardware For Real Time Computing [51:49]
Patrick:
My history with this was back at Bell Labs. We were monitoring numbers of phone switches in real time. And in order to do that, it was required, this was quite a long time ago. This was the era of the 5E, that should explain how old this was. They were quite new. It was one of the first digital phone switches. But it required massive amounts of customized hardware to monitor these switches.
Today, I do not see the investment in that type of customized hardware for real-time computing as much. It is kind of uncommon to see really, really specifically focused hardware for real time computing like we used to see quite a lot. You would see real-time Linuxes. You would see real time that were made for all sorts of different RTC clocks and RTC controller cards and things like that, to ensure absolute peak timing and synchronization.
We see some of that today, but not nearly what we did maybe 20 years ago. I think that is either because the performance of machines has gotten so much better and we are close to an okay level, or our standards have slipped a little bit. Or maybe a merger of both, I am not sure.
Zane:
That is great. Thanks, Patrick. We do have a question from Sylvia. Have the programming languages and interfaces changed over time?
Changes Over Time of Programming Languages and Interfaces [53:39]
Greg:
I think It is a great question. Yes. When I was first getting into this, Gary, you may attest to this, we actually had to help a lot of scientists and researchers upgrade to Fortran 77, because they were running predecessors. I forget what Fortran it was before 77, 60-something. That was the year by the way that the spec was created.
Yes, it has definitely changed. We have seen a lot of that. Now, on parallelization, we have seen NPI really taken off. We have talked a little bit about the different forms of NPI. The specification has definitely evolved over time. That specification has created a lot of compatibility. So even though we have many different NPIs, coding for any of those NPIs should be the same no matter what, if you are adhering to the strict specification. There are some alternative or optional pieces of the specification, but as long as you adhere to the strict specification, it should be NPI compatible no matter what.
Now, every programming language is going to have its own aspect or interface to NPI. You could have, for example, C bindings, Fortran bindings, Fortran 90 bindings. More recently, we have seen a lot of interest in higher level programming languages like Python. Python has absolutely taken off in scientific computing, to the point where I actually think it is probably one of the most dominant programming languages nowadays. I am curious from others: are you seeing the same thing? And what is that like in industry-specific areas like life sciences? Is it similar there too?
Python and Scientific Computing [55:35]
John:
Our software stack is almost entirely Conda at this point. There are a few applications in there, Bowtie, same tools, that kind of stuff that is not in Python. But for the most part, at least 70 or 80% of everything is done in Python. Maybe with a mix of R in that too.
Greg:
Glen, are you seeing the same thing in your area of life sciences?
Glen:
Yes. Things have evolved from Perl to Python. Python has really done a fantastic job in adding libraries. Then of course, Conda to support the researchers. It is, I think, the most popular language. There is a lot of R in there. From folks who come from a statistical side, they seem to like R a lot. But Python is, this is just my opinion, but if you are trying to do a lot of other things than just statistics, Python is much better. You are not bringing up a web server in R. We see Python doing a lot in the life sciences.
The folks that are using C are the folks writing the algorithms, whether it’s for DNA alignment or drug docking or on the GPUs, trying to get their stuff onto the GPUs with CUDA. Those are the folks using the low level languages. Everybody else is a little bit higher up at the Python level. There seems to be a little bit of a push to, I can't remember the name of the project, but actually getting some actual typing, static typing as opposed to just dynamic typing with Python for some things in research. I am actually interested to follow that. Yeah, that is what the trend has been.
Greg:
It is funny for me to even consider and think that HPC has changed so much over the years. We are not even focused solely on compiled and purely efficient languages that we are running across a whole bunch of different languages. I think there's even some Rust that I have heard of, Patrick, I think you've mentioned Rust as well, coming into HPC.
Rust And HPC [58:07]
Patrick:
There is Rust. Also there's still some Ruby creeping around in the EDA world and moving in there. But primarily these days it is Python and Rust and again, some Ruby and of course, sadly Perl still exists. The Swiss army knife.
Zane:
You are making Greg cry.
Greg:
Okay. Are we done with the webinar yet? Come on.
Patrick:
No, I love the Swiss army chainsaw, but I also hate it. If you have ever walked into a project where you have had somebody whose been in their own Perl world for years and years and years just designing it and designing it and designing it, and just adding on. Then you try to take over that project and figure out what was actually done with this stack of Perl scripts, it can be fun.
Zane:
Then you rewrite in Bash in like 10 lines.
Patrick:
And make you hate Perl.
Zane:
Yes. We have a new question from Dave. I have seen so much SciPi and NumPy usage at scale these days. There are a lot of compiled time optimizations to be considered for those libraries alone, Intel compiler and KL.
Scientific Libraries For Higher Level Programming Languages [1:00:20]
Greg:
Yes, there are a lot of scientific libraries that are also becoming more and more available for Python and other higher level programming languages like SciPi, NumPy and others. I have not seen, and I am curious if others have seen, architecture optimizing math libraries like MKL, Intel MKL. Are there interfaces for Python for these? Or are we moving away for the specific need of high performance, absolute tuning, profiling and building the fastest programs ever in exchange for more general purpose HPC applications? I am kind of curious. I know we are at the top and I am really curious about people's take on that. Maybe that is where we leave off for our next round table.
John:
It has been at least seven or eight years since I have built a Blaze library to do specific optimizations. I think at some point the hardware got fast enough. People, certainly at the level I have to deal with, do not care about getting the last five or 10% out of anything. The next generation of processors will solve that problem for us. They focus more on writing a better algorithm in general, rather than trying to get a little bit more compiler optimization out of it.
Glen:
When Intel came out with their Intel Python and then MKL and Blaze and Boost and things like that optimized underneath, we were really excited to use it, but it just did not seem to play well with the typical Conda, PIP tools. When we tried to install it, it just kept blowing up on us. We gave up on it because we thought we could get a lot of good performance gains from it because everyone was using Python, but it came down to just usability.
Or just trying to introduce it into the already mature Python environment everyone was familiar with. We did just try to use these libraries and things like that; it just did not work. We eventually dropped it. Other than that, we have not built or compiled anything specific for tuning or a library in it in a really long time.
John:
My most frequently used C flag is M2 and Generic, just because I want it to run everywhere. I do not care if it’s fast, I just want it to work everywhere I am going to run it.
Patrick:
See, that is what it really comes down to, is that you can make snowflakes that work really, really, really well one time. Making something resilient that can run everywhere all the time and reduce your issues is so important to HPC in my opinion. If you are constantly retuning, retuning, retuning for different architectures and different last 2% of optimization, then how much time are you losing to get that 2% when you could just throw more hardware at it, throw more cores, throw more memory or throw a cloud at it? That is the other aspect, is now that you can fairly easily scale up. It is easier to do that than it is to optimize just down to that little tiny little bit at the end.
Zane:
Thank you for that, Patrick. We have one more question before we wrap up. Are there analogs to MKL for things like the ARM64 architecture?
ARM Clusters [1:04:24]
Greg:
I do not have an answer specifically for the question. I would assume if there is not already, that they are coming soon. I just wanted to say hey to Dave, It has been a long time since we chatted. It is great to see you here.
Zane:
And thanks for all the questions, Dave.
Greg:
Does anybody else know if there are optimized versions of math libraries for ARM? Dave, we are going to probably have to reach out to some people that are running some ARM clusters. I know Sandia has some pretty good size ARM clusters as well as maybe some ARM people. We will hopefully have an answer for that either in the chat or maybe for our next webinar.
Zane:
I am taking notes. Well, guys, we are a little bit over and we do appreciate your time. We know you guys are busy and thank you for joining us again. Krishna, it was nice to meet you. Gary, John, Patrick Glen, Forrest as always. Greg, thank you very much. Guys, go out, like, subscribe and we will see you next week. Thank you very much.
Transcript
good morning good afternoon and good evening wherever you are we welcome you back to another ciq round table so we had one of these week before last and it went over really well so we invited a group back to continue that conversation and kind of see what else we can we can get from people who are in the industry and actually have real world use cases and talking points so we could go ahead and bring in the panel oh god it's greg well welcome back everyone we're missing we've added some new ones this week too i think we're missing one as going to join if
not we'll go ahead and get started so i think last week every everybody got to meet patrick and john and gary and greg as always so i think we'll start with a new guy in the room if you would go ahead and introduce yourself and [Music] way with greg kudzur and gary jung both of them here on the panel i before this i was employed at torres berkeley lab at uc berkeley craig hired me into that position at berkeley and gary was my math manager and guide and mentor so yeah i know the panel very closely um i have worked with patrick i have xj
on on slack i've i learned a lot from patrick i know him uh glenn good friend of mine um yeah i'm glad to be here uh so i i i'm about 10 months old at kla corporation this organization we make wafer inspection devices um so in semiconductor industry it's very crazy right now with all the chip shortages uh the demand for our product is really high and it's a little bit crazy and we are having to stretch ourselves in terms of our hpc designs and htc technologies we are we are not adopting uh to meet this growing crazy demand um yeah this topic is really
interesting i'd like to be here very nice and welcome so i'm gonna go around this the screen real quick i just entered yourselves very quickly again uh i think because you've entered yourself last night to go into a lot of detail but i'll start with you patrick because you're at the top of my screen here uh howdy uh well my name is patrick roberts um i currently work with skyworks um skyworks we design all sorts of different types of uh chips for all sorts of different things so uh sorry that's rather vague but we're in all sorts of rf chips analog chips digital chips uh
View full transcriptHide full transcript
big and small so um been around for quite a while um they used to be rockwell long long ago um so some of you may for me may be familiar with that it depends on depends on your age i suppose but uh i've been in hpc specifically for a little over 20 years i've been in the it industry for 29 29 or so years so that's about it it goes fast does it patrick it does it does indeed so next we have john aka all grizznock uh john hanks been doing this for a long time i don't remember what i said last time when i
got introduced so i hope i don't contradict myself this time um i i was thinking about the topic today and i think maybe i'm inherently unqualified to answer this question or even try to answer this question because in in pondering it i i think all i do is manage workstations now they just happen to look a lot like a cluster because of the way i manage them so i don't know that i even do any hpc anymore no i think that's great that's something that's going to be interesting to kind of see i mean what is hpc to everyone else so i'll move next kind
of skip over christian because we had you greg you're in the middle hi everybody i i feel like who is in the brady bunch uh in the middle was that alice jan no i'm kidding thanks um but uh hi everybody uh i've been doing hpc pretty much my whole career started with a degree in biochemistry and then failed to continue getting my phd because i got more interested in computing and linux and open source than than than doing the science so i've been really focused on hpc and scientific computing ever since and as christian mentioned um i i've worked with krishna and gary for ages
um and at berkeley lab so uh it's great to have great to have this middle line across the board is is is the berkeley is the berkeley crew here i think very nice we'll just call it the berkeley line there so next person from berkeley gary gary hi hi my name is gary chung i'm i manage the institutional hpc computing at berkeley lab greg and i started the program back in 2003 and it's been going great since i also have a dual appointment with uc berkeley so i managed the hpc effort for the institutional uh institution down at uc berkeley also um yeah and uh
krishna and greg and i we are way back doing this very nice glenn you're up next can you guys hear me now we got you yay uh glenn you gotta you gotta you gotta show us your shirt today too by the way [Laughter] uh been at ciq for 16 days i'm the uh vp of freaking sweet ideas the iq and uh my background uh was a doctorate in immunology of microbiology postdoc in hiv research and then um left academia and fell into hpc with my uh second postdoc at the san diego supercomputer center and uh from then on i was like in 99 from then
on it's been you know bioinformatics and uh hpc has been has been my livelihood so uh incredible and since about 2000 i've known greg i think so i blame him it's an easy target right yeah he's in the middle exactly hey everyone my name is forest bert i'm a high performance computing systems engineer here at ciq i've been with ciq for about a year before that i was the student hpc system administrator at boise state university while i was getting my bachelor's in computer science i kind of got my start in the cloud and in computing and stuff like that um you know in the
early 2010s messing around with um you know the tech that i had available and uh you know cloud gaming servers and stuff like that i'm real pleased these days to work with a little bit bigger systems um but yeah i love hpc got super interested in it while i was in college and uh yeah i'm really pleased to have been working at ciq and to uh be working on everything that we do with fuzzball and that type of stuff so great to be here thank you forest so kind of jumping back into the question of of kind of what john started to allude to a
little bit and i think that the the definition of hpc is going to be a little bit different to everyone here so kind of like to go around and ask what does hpc really mean to each of you and just kind of tell us what that means tell us a little bit about maybe where you got to that point and how you got there and start off with patrick and keep picking on me adam you're right next to me uh i would say what does hpc mean to me so hpc to me is any service which is designed specifically to perform a set of complex
tasks which i know that is a very broad definition but it's very um it's very true because i've dealt with hpc in all of its various forms from um from craze to to working on loose grid systems to working on top uh uh close coupled very tight systems with infiniband and mirinet and you know just all sorts of different interconnects of all kinds and it really just depends on what tasks you're trying to accomplish what you need in order to meet that task out um you know it kind of there's a blending of what is it there's there's hpc for high performance computing and then
there's now high throughput computing and now there's like all these other sub variants but in my mind it's it's really it's all hpc it's just with different tasks in mind to accomplish and and and different hardware and software solutions to accomplish those tasks so for me it's just kind of one really big bucket of uh designing clusters and computers to meet tasks needs so that's great thank you patrick and john you kind of started down the road of talking about desktops and that's kind of what you would deal with today i know they may have been a little bit tongue-in-cheek but in reality that can
kind of be something that that really is the case so tell me about your your definition of hpc for for me hbc the primary thing there is high performance and for it to be hpc it needs to be high performance and that's why a lot of what i do now is not hpc because i got people getting you know this massive pile of gpu hardware and and memory and cores and then sitting there typing in a jupiter notebook on that system it's like the very as far away from high performance as you could ever hope to get you know somebody checking out a seven thousand
dollar gpu to type and hit enter in a command line prompt um so i i think it's the hpc high throughput computing those are really the two main designations and then for me it's getting even further away from high throughput computing where yeah we run some embarrassing and parallel workloads but the most important thing in our environment is on demand for people interacting so we're we're pushing that long tail even further out by pulling in people that don't really need a cluster they just need a well-managed desktop workstation that has a lot of memory very nice thank you very much krishna how do you define
hpc i i think it is changing it is my audio better more um you're good okay um i think the definition of hpc is changing 2005 when i started my first job in us it was at san diego super computing center and the first machine i worked on was from ibm it's called data star it has a very specialized interconnect uh very specialized work servers all put together but that's not the high-end machine back then a san diego supercomputing center had a good size machine data star that's how the hpc mission used to be and the applications that run on used to run on that
application was very optimized tuned for the architecture that ibm architecture that we used but these days i think it's not the case anymore a lot of technologies have commoditized and interconnect infiniband as it has is is treated as a commodity a lot of servers that we use on even the top machines the top 500 machines our commodity servers from supermicro and things like that i think it is morphing the definition for hpc is becoming a lot more commoditized and not specialized machines anymore the employer who i work for right now we build stream computers all that we have is a lot of data streaming at
us um wafer inspections right so we are we have really high high definition cameras uh looking at this nano scale uh layer width uh transistor width and things like that so it's really high volume data just fire hose of data being shoved at us the biggest machine that my uh my employer built is just three racks not hundreds of racks or even tens of taxes just three racks that's the biggest that we have so and it's all commodity hardware commodity we we don't have a specialized server design team at kla we don't have a specialized architecture team but really we all take white boxes super
micro band commodity and just try to build that this is i would consider what we are doing at kla is hpc because the data volumes that we are dealing with um if that's the definition it does come under hpc but the hardware wise it's all off the shelf so it is changing hpc definition is surely changing is what i would say oh that's very interesting i mean talking about commodity hardware versus very specialized hardware that it used to be so i think like you said things don't think there's one way i don't think that's inherently a bad thing though i really think that commoditization of
the hardware has brought hpc into the hands of so many more people that it's actually a very good thing that you know you can build a super equivalent with off-the-shelf parts i mean i i did it for for a a very very small um but one could argue quite important x86 design company where they didn't have the money to go out and and and really invest in a massive super to to kind of solve their complex problems uh but instead i was able to put together desktop machines literally on a folded piece of metal and i would go to fry's rest in peace and purchase
off-the-shelf hardware and literally take it back and have one of my guys put it on a folded piece of aluminum and bolt the the the bits to the they'll fold a piece of lumen and have bread racks from sam's club with like literally 44 machines per rack in an open air data center because it was designed originally for uh for uh mainframes so it was you know designed to to kind of have an open-air concept uh and you know roll in the bread racks of 44 machines and we could do it and add thousands of cars at a fraction of the price that it would
take for us to go to ibm or super micro even or dell or anybody to have it done and so i mean just the ability to do that um it brings you know it allowed it allowed that small little hundred person company to compete with the likes of amd and intel and not necessarily effectively compete um as they got bought by intel last year but um they were able to compete so i mean just that that that alone you know that commoditization of hardware it's not a bad thing in my opinion yeah it's really cool to work on those really highly specialized systems um but
it also is it's amazing what bringing that capacity for computing to the general populace can really do well before i go to greg for his definition of it i i coming from an enterprise background i've watched people build commodity infrastructure to solve enterprise problems that probably should have been hpc in the past and it's interesting to see kind of that merging of the two worlds where you see hpc becoming more commodity hardware and maybe those workloads are moving from just general enterprise into an hpc environment where they actually get handled in a much more efficient fashion so i i've been excited to kind of see
that merging but now i'll go to you greg what's your definition of hpc well before i even mention that i just want to respond to patrick's point i completely agree we've seen this happen so many times in research organizations especially governments in academia where in many cases people you know that have never done system administration before are even buying clusters and you know they're not managing them very effectively but they're getting research done and because of small budgets they're actually able to do uh very good research on that at some point typically they'll graduate and and be able to then to have uh you know
proper support for those systems so it's not you know a grad student or or even an undergrad or a postdoc who lost a bet and now has to go maintain that system uh in many cases then they can actually reach out to somebody like you know gary's group um at berkeley and and lbl to actually solve that problem and actually manage it appropriately but but it hitting commodity has lowered that barrier of entry so high that it's been it's been fantastic now to get to the question you asked me on zayn um you know i've been doing this for long enough to have a very
strong opinion of what high performance computing is or rather was uh it was it was tightly coupled uh parallel processing at massive scale and i remember times where unless you can prove that your application could run effectively at massive scale you didn't get time uh you were actually uh you know you were down revved in terms of um uh how much time you can get and what priority you get to run in unless you can demonstrate that level of efficiency now we used we called it the long tail for a long time which was all of the scientific applications and work workloads that are running
on these systems that are not that massively traditional high performance computing uh profile and what we've seen over time at least what i've seen over time is that that long tail has just continued to grow grisnog you kind of alluded to that as well that the long tail has now grown so far out that the long tail is now becoming the dominant form of workload on many of these systems and i've had many many arguments over drinks at super computing on what is high performance computing and um and is these sorts of applications actually high performance computing it took me a long time to actually
um to get to this point uh in in my maturity uh as a person to actually say that high performance computing is now more than just those massive tightly coupled applications it's anything that's going to spin or peg some aspect of that hardware at 100 percent and then literally kind of bottleneck based on that that subsystem of your hardware so whether that's cpu whether that's memory io whether that's bandwidth you're going to peg something on your system and that's kind of where i am in terms of the definition of high performance computing it could be a single thread hitting one type one core or even
a a hyper thread on your on your cpu but it's it's pegging something dang it somewhere that's that's where i draw the line i like it efficient utilization of resource that's all i'm hearing i'm not sure if it's efficient in many cases but it's definitely utilization pegging it resource you know whatever gary your definition you've got a fairly large environment to deal with so yeah yeah um you know we uh it's interesting uh when when greg talks about that it brings back a lot of memories i i uh it was exactly the way he said we're looking at really big um usage of the system
but you know we we do take a a kind of a simplistic look at view of like what is hpc we work with a lot of researchers team large team science and then also individual researchers and a lot of individual researchers may have a very powerful workstation they may have a laptop and then at some point you know what they're doing data acquisition or simulation they're going to outgrow it and so they have to move to something and then so what is that something and then for us that that's where we define the hpc we want to move them on to um like you know
instead of coming up with one offer all these people then that's that's our opportunity to move them into an hpc infrastructure and then you know i i think everybody's kind of covered that but i there's enough convergence in the in the field so that now when you describe an hpc infrastructure everybody knows exactly what you're talking about so one thing i just want to follow on that gary said speaking of memories um we did see a lot of people that started off on a laptop then went to a powerful workstation and then we initially when we first started the hpc project at berkeley it was
around this concept of something called mid-range computing and mid-range computing was we we defined as everything between a workstation and a giant facility like nersk and everything in between uh nursk at that point was um limiting people unless they can prove you know a reasonable amount of performance and and efficiency in terms of running on their system they didn't want applications running on that system unless they were efficient on that system because it was such a big expensive resource so one of the one of the big motivations that we had for this group was uh this idea of mid-range and and how do we solve
these mid-range problems and that to me is now turned into what i kind of call sweet spot computing this is where most people are are focused on this is this is for example in high performance computing this is the vendor's bread and butter this is where they're selling the most systems so and that's why we're talking about turnkey hpc honestly because it comes right back to that that's the problem set that we actually really need a solution for yeah that was a that was a huge gap i i think essentially when we did that that early you know we helped to define what is known
as the mid-range computing segment within the doe and i think that actually even articulates how hpc was changing even then gary without us even really kind of coining it or thinking about the the nomenclature we were changing what was hpc even then it started off being those giant systems and then the mid it was moving towards mid-range and it's been still kind of following in that trend in that direction ever since it's very interesting thank you very much gary next we'll move to glenn you have a interesting perspective on this so what is your definition of hpc so so when i started uh in this
field hbc was you know tightly coupled pretty large scale machines right marinette low latency interconnects but when the life sciences started adopting the beowulf architecture it rarely had um those little latency interconnects because it didn't really need it it just needed so they were coining it high throughput computing right on on ethernet and so i read it to people oh it's not really hpc it's just high throughput but you know in my head i was like same architecture just it doesn't have your low latency interconnector like we're doing and it's a different you know type of parallelism um but i still thought that it has
to be something like hpc and then in my head it kind of evolved to like so how do i you know lump those two together so that and then that became research computing so research computing was you know all these folks whether it was uh molecular dynamics or or just you know dna alignment embarrassing the parallel or tightly coupled it was all research computing but then the enterprise was doing it so you know so so now what do i call it um and then to john's point you know i've worked on machines with you know all these gpus they're running a jupyter notebook but you
better have that rapids library and that python tuned optimized really well so they get they get the most out of it and so that if my last employer when we were working on my title which was not vp of freaking sweet ideas unfortunately uh we talked about why not scientific computing you know you're trying to get science done really whether it's you know johnson and johnson trying to figure out you know how to design their next shampoo bottle or if it's you know covered you know um kind of genome analysis so i've come to call everything that we've just talked about scientific computing but if
you hadn't if you had to like nail me down to hpc i'd be thinking of systems that you know try to look like the top 500 right tightly coupled uh the scale of the top 500 um anything you know smaller than that is you know still hpc um i'm not sure my mid-range hpc um redline is but um that's that's how i kind of view the world these days thank you glenn so kind of between what greg talked about and what you talked about i used to be on a mid-range team and mid-range to us just met the solaris team and one of our engineering
stations yeah yeah exactly one of our guys called himself the vp of blinky lights because we had the data center so that's not quite as cool as sweet ideas but pretty cool forest what is your definition of hpc i would i would define hpc kind of surrounding um the efficiency that we're trying to achieve with hpc we have you know hypothetically a single commodity server that we're trying to you know do some traditional hpc workload on whether that be you know simulation or you know just to roll them both into one thing is some type of embarrassingly parallel high throughput type of job you know
if you only have one server you're not going to be able to get the performance that you would need out of that in order to actually um you know efficiently iterate on the work that you're trying to do um so to me hpc comes down to being able to do these really really complex things in an efficient way um regardless of what the system underlying it ends up being whether that's you know a bunch of ec2 instances that's you know a traditional tightly coupled cluster whether that's a cluster that's more meant to provide as many gpus as possible just to you know do mass model
training that type of thing um i think ultimately it comes down less to the comprisement of the system itself but are you building the system to gain performance beyond what you would get with you know the base component of this system in that case i would say you're doing hpc i would also touch on kind of the long tail concept um because uh kind of where i very very first entered this was more so in those tightly coupled jobs but especially at my last institution i started to see a lot of researchers coming out of very what i felt were non-traditional places for hbc um
you know we weren't just having um you know genomics people or simulations people that type of stuff we also had you know education or people from um you know the education department that were trying to do mass data science crunching for their research we had people kind of on the high throughput side we had people that were taking lidar data and doing um huge amounts of analysis on that kind of in a high-throughput way so i definitely think that in addition to it kind of being about efficiency we're definitely seeing you know kind of enterprise and some of these other use cases realizing that they
are trying to do hpc without realizing that they're doing it um so we kind of see uh you know maybe some of the muddying of the waters here as some of these non-traditional uh entities kind of enter the hpc sphere absolutely and that kind of leads into one of my next questions for us and thank you for that is i think we all alluded to it at some point but have the hpc workloads changed over the years and we talked about the infrastructure getting better the machines getting better i mean obviously they're getting far far more and more dense and being able to be more
capable but have the workloads changed over the years and i will throw that out to anybody who wants to take it 100 absolutely the workloads have changed i mean you've got some of the traditional workloads that are still either that are still around in mpi mpi-based ones we touched on this i can't remember if it was in one of our dry runs or or in the actual uh uh meeting last week but we we definitely touched base on the you know who's still using mpi what are you you know what are your percentages of your workloads of different types and that sort of thing but
it was it was definitely uh you know eye-opening i i kind of expected it in some ways but it still was very informative that like that the mixed workload is really here to stay um even in very specific like industries the amount of workloads that yours you know singular cluster or clusters are supposed to to be able to support is massive you're supposed to be able to support you know gpu-based uh you know ai training versus or ml training versus you know being able to to to run you know ultra fast single threaded performance um you know or highly parallelized jobs you're supposed to be
able to run it all in one and you know all of those require different really pain points to overcome in building hpc clusters but i think we're kind of expected to you know as designers and administrators and of those environments we're supposed to be able to provide for all of those in any kind of one size fits all and it just it's difficult to accomplish i i think to some extent um the commoditization of the compute part of this has made maybe hpc simple to do on the compute side but where the actual interesting stuff being done now and probably more so in the future
is on the storage side storage is is actually becoming the critical thing about everything we do regardless was hpci throughput whatever i mean at least for life sciences there a lot of on-prem clusters and uh on-prem infrastructure exists today because you just simply can't move the data sets around like the compute has to be brought to the storage so the compute's more or less secondary to just getting all your data in one place where you can work on it and not only that the way the compute is being orchestrated or designed is quite different in the in the enterprise enterprise is dealing i know of
use cases where they are dealing with um data pipelines at the rate of multiple tens of gigabytes per second constantly coming the that's the data volume that they're dealing with and the workflows and applications they're running they do not have anything mpi no mpi at all nobody in that organization knows anything about mpi so it's uh it's not yeah mpi is not the only hp the data volumes are a lot higher than many organizations deal with but these organizations don't have anything mpi lack of lack of mpi in the enterprise um pushes me into an old man rant moment about a negative connotation of hpc
and that had in the early days of hbc i remember back you know 20 or 30 years ago when i was getting started in this there was an elitism aspect to hpc you know you had to be able to afford a cra to do it it was a very elite club and not everybody got to join that club and that attitude carrying forward over the following 20 years from from when super computing got started led to us having to deal with things today like hadoop and spark and kubernetes and all this just nonsense complicated extra complexity poor performing you know because that's the stuff that
came out of the enterprise when at any point during that we could have just said hey enterprise people how about you help us write a better mpi and let's do this right uh you know efficiently and simply with a simple scheduler i don't think it was that it wasn't offered i think that it wasn't necessarily accepted in amongst industry because so many different corporations were trying to capitalize on that market or well here let me enter my proprietary standard into this this ring let's let's go ahead and and they didn't want to do it open source either generally they wanted to really make it very
tight and and you know hey well now that you've entered into our ecosystem you're stuck and that's you know that's kind of the way that came about i think is it and then of course you had the the competition as you said where now we've got coop and now we've got you know all of these other uh competing not exactly optimal frameworks to put it lightly uh that to do a lot of the same workload that yeah you're right could be could have been much better implemented in hpc so let me ask this pack and john so from an enterprise perspective if i'm looking at
this the way that most of those other languages and other tools came out was to be simple and fast and if i look at mpi it's probably a little more complicated to program for correctly and to get it as efficient as it should be right so was it something that from an enterprise perspective people didn't want to take the time to learn how to code correctly for mpi it was just easier to use some other framework i'm going to throw on java it was a major barrier to entry so because of that barrier to entry whether it be i didn't have the hardware to have
that high speed memory interconnect between the nodes or i didn't have you know ca i'm a you know undergrad i can't afford to get time on a on a super at the lab i couldn't you know or whatever it whatever it was you know uh there's always a barrier to entry or if it's well if it's intel mpi do i have to buy into the entire intel ecosystem do i have to go buy an intel compiler back in the day does that mean that now that i'm running on my amd chips is it going to be slower because intel spiked the compiler i mean you
know whatever it is you know it just it depends upon the the environment but there was always a barrier to entry um and then of course then you had open mpi so now we've got competing standards uh so i i'm gonna i'm gonna push back a little bit because um i think it's very use case specific so the use cases and the applications that we were solving in high performance computing were very different than what enterprise was really even looking at solving now the map reduce um kind of focus there i do i do agree there we could have done better in terms of cross-pollinating
between hadoop and then spark and then then hpc mpi and piio and other things and we're starting to see that also happening in in ai starting to right this is a few years old now at this point but things like how do we do training how do we optimize training how do we start parallelizing training if anybody remembers some of the early training frameworks uh it wasn't running mpi they they created something from scratch almost like mpi didn't exist and they solved very you know it was it was a duplication of effort and to the point patrick i think you made this um it may
have been grisnog but if we would have come together right at the beginning in a better way more open way and i think showed more interest in collaboration i think we would have seen um very different outcomes uh and and but this goes back to i mean i joke about this uh you know in in all of the organizations that i've seen and worked with that have hpc focused people traditional hpc focused people there's i t and then there's hpc even if hpc is under i.t they don't talk they don't work together they don't leverage each other because the technology is both so different but
they're also they're solving such different problems but there's also just this legacy mindset of on both sides it's not until fairly recently in my mind containers was the pandora's box that that really started blurring this it wasn't really until containers did did anyone in hpc even care what the heck was happening in enterprise we took off at that point and on the enterprise side it wasn't until just recently did they start looking at hpc saying how do we do training faster how do we do better analytics how do we get more insights from our data and that to me was the crossover point and now
we need to actually all come together and realize we're solving very similar problems and we need to stop innovating on our own and start back start actually working together and solving problems together we need to be a big happy family again need to be yes seemed like you were you were going to what do you mean what do you mean again um all right i just kind of want to rip off that a little bit um of a side jaunt but back to the topic of how hbc is kind of changing um something that i really see is uh the hardware side of hpc changing
it seems like for a long time you know we had you know infiniband cpus gpus and that was you know kind of the basic compromi are you know the basic components of a cluster as far as you know the compute part we start to see now that there's a lot of companies that are starting to invest in creating new solutions around some of these very specific use cases so you know now we see for ai you know we're only 10 years off of alex uh krishevsky i think i pronounced his right name or his name right there um you know training alex net on two
gtx 580s and now you know 10 years later we already have companies that are trying to replace the gpu for ai training because there's better ways that that can be done than just on gpus um so some way that i really see it changing is on the hardware level i mean we've got um new ways of training ai new uh much faster interconnects than we've ever had available before new storage technologies it seems to me like we're really in kind of a renaissance of hpc hardware um you know the uh the use of pcie as a broader interconnect than just to put you know accelerators
into a cluster and stuff like that uh we're really it seems like kind of in a renaissance as far as hpc hardware goes um and our clusters are really kind of starting to change from just the basic cpu interconnect gpu type of thing with fpgas specialized chips what have you of course i'm glad you brought that up because that kind of reminded me of a question i had for gary last time is he made the statement that his gcp usage has gone up by 3x over the last couple of years a few years now what is driving that and is that something that is more
hardware specific since it is difficult to just go purchase whatever you need because it changes all the time is that kind of what you're seeing is that what's driving that yeah you know i for a lot of people they're working with workflows if somebody's um wants to put together or string together uh an infrastructure with a web server and a database and do something and move some data is pretty easy you know the term i like to use is is frictionless computing because they're able to kind of do it themselves and without having to say wait for a network drop in a data center and
then for somebody to set up the machine and all that so it's a lot faster that that's it and then uh probably the other thing uh is this really easy to prototype things and then the new tools like automl and sagemaker make it really easy for people to get started using machine learning and so um so it's interesting we'll see we'll onboard people onto the cloud we'll kind of like prototype it using auto ml or something like that you know get going and then then sometimes what happens is that they'll figure it's find out that it costs a lot of money and then we'll go
back and now that they see how how efficient it is and how this helps their science uh we'll go back and um transition on onto our on-prem hpc um using uh you know maybe slightly different tools interesting but yeah i was just curious because that's it's interesting to see how that's taking place and i talked to a lot of different customers who are hpc that are trying to like you said try things out in a cloud before they actually make the investment to see if it's gonna to work for their needs or do what they expected to do so thank you i think we have
a question from mr ignite so i before we pop it up i can go ahead and start reading it is nobody's really talked about the vector processing as a part of hpc build out so the the example that's given is a medium scale iot collection and analysis of billions of devices generate tens of bytes per second driving anomaly detection pattern recognition and other applications uh interesting interesting topic and if you guys have any thoughts on that i'd be curious to hear if anybody's ever done anything on a cluster that was in real time and the the closest i've ever got was a group that wanted
to produce a weather forecast every day and so they had data arrive and then they had a deadline to get stuff out but i've i don't i don't recall ever having anybody want to do real-time stream processing as part of a cluster workload um actually krishna and greg and i could probably have an example of where we're currently building a system for a gamma ray detector and this uh the system is going to be the uh the first level trigger of the data coming off the detector and um it's going to be a kubernetes cluster we're going to run different containers depending on what's going
to be analyzed but uh we we are using instead of fpga hardware they figured it's going to be a lot more flexible do with the cluster so that's currently uh being put together engineered and put together right now so to kind of and i kind of figured you know gary and krishna would have some examples here but um yeah there there were multiple different uh examples of people doing detector research and uh they basically you can you can think of a detector as basically you know a very kind of simple idea of one is you know you have some sort of reaction happening within a
a spherical enclosure and then you've got a whole bunch of detectors all the way around the spherical enclosure and it's it's monitoring all of the different types of rays and and what is happening at the center of the sphere and all of these different detectors are uh creating streams of data then these streams of data have to be computed in parallel in real time so you can see and visualize what's happening within that experiment and sometimes these are actually in real time but in many cases it's near real time which basically means now you actually have to um look at it and then decide do
you want to rerun the experiment do you have good data and then feed that back into your into your process flow but it needs to be either real time or near near real time for that and as gary just mentioned if you're running it through kubernetes it's an extraordinarily different architecture than what we've been doing in high performance computing and without even knowing the specifics on that i would say probably it's a bunch of services that are running and doing the equivalent of like an inference on each one of those data streams coming in and then computing on those via gpus so this is like
a performance critical service that's running and kubernetes is a is a good architecture for managing services um i i'd actually i i i'm curious on the efficiency that you're getting because i've actually had really bad results with regards to kubernetes and just to go on a quick tangent real quick uh we actually were talking to some people that were doing something along these lines with kubernetes but because kubernetes is so big and bloated uh it was usually utilizing about 15 to 30 percent of the underlying hardware resources uh just to run kubernetes so now let's scale this up imagine buying a 5 000 node cluster
or even a 500 node cluster and 100 100 nodes out of your 500 nodes is just running your your your scheduling system your your orchestration system right and you're you're basically losing 20 of your system to run just for kubernetes so at a small scale i think kubernetes would do it but as soon as we start scaling bigger we've got to come up with better ways of doing that yeah another another approach i have seen in the industry for these real-time stream computers is a traditional way which is so if you look at a ddn appliance a storage appliance um they usually come in pace
and they come with some fault tolerant resilient designs using basic technologies like ha proxy korosync such that services can flip from one controller to the other controller automatically if if there is a failure in top controller such services can luster services can move on to the bottom controller that's the basic components that you nee you would anyone you would use for fault tolerance um i'm sure this even kubernetes is kind of built on some basic components like that um in the industry i have noticed if you build a real-time streaming computer where your machine goes into a production pipeline and they need real-time feedback of
the quality of the product being being manufactured and if this computer is down and that quality is not being generated depending upon the value of the product you are making it can cost multiple hundred thousands of dollars per day of production loss so that's very costly so this real-time computer giving the feedback on the quality of the product is a very critical component and when you're making such a designing such a computer fault tolerant designs having backup components so critical and things falling back automatically from one server to other server is such a critical thing in your design if you do not have it you
you are costing hundreds of cases of losses to your company to your product um that is something that i have i'm noticing heavily and not everybody in the enterprise is adopting kubernetes they they're just not trying to accomplish that using the ha proxy chorusing services which have existed for a long time uh and trying to put together everything um like we have redundant switches redundant network cables run from uh servers to switches and how services fail over from one switch to the other switch if there is a switch failure or a network cable failure those are those are another tangent an extension to hpc design
that i am seeing in the enterprise and that is very critical and needed when you are making a building a real-time computer for these use cases so from my past iot experience most of what we were trying to do was do as much analyzing of data at the edge as we could so that we weren't sending massive amounts of data back to be processed or analyzed you were looking for that anomaly and then whatever data was driving that is what we were sending back so from a real-time usage i feel like just because we were trying not to send a lot of data that may
have been irrelevant we were doing a lot of pre-processing at the edge and then only moving back what was necessary for the next type of action or the next time event that needed to be taking place but hey that's an interesting question when you're talking about massive data like in greg's example it's not really an edge's case because everything is right there together you're trying to do it in the place that it's taking place from an iot perspective when you have billions of sensors out that amount of data streaming back could be something that would be overwhelming and a lot of that at the end
of the day just gets dropped on the floor and you never even look at it so a lot of different ways to approach it when you look at iot it kind of depends on what you're trying to accomplish in my opinion my history with this was uh back at bell labs they we were monitoring uh numbers of phone switches in real time and in order to do that it required i mean this was it's quite a long time ago this was the fi the error of the 5e which that that should kind of explain how old this was but um and they were quite new
it was one of the first digital phone switches but it was uh required massive amounts of customized hardware to monitor these switches and today i don't see the investment in that type of customized hardware for real-time computing as much i mean it's it's kind of common or i'm sorry kind of uncommon to see um you know really really specifically focused uh hardware for real-time computing like we used to see quite a lot i mean you'd see you know real-time linuxes you'd see real-time um you know that were made for all sorts of different rtc clocks and rt rtc controller cards and things like that to
ensure you know absolute you know peak timing and synchronization and we see some of that today but not nearly that we did maybe 20 years ago so i think there's either because the performance of machines has gotten so much better and we're close to a okay level or our standards have slipped a little bit um i think or maybe kind of maybe a merger of both i'm not sure but that's great thanks patrick we do have a question from sylvie and it's have the programming languages and interfaces changed over time i think it's a great question um there's yes um when i was first getting into this uh gary you you may you may attest to this um we actually had to help a lot of scientists and researchers uh upgrade to fortran 77.
um upgrade two fortran 77. uh because they were running predecessors i forget what fortran was before 77 60 something and that was the year by the way um that the spec was was created uh and yes it has definitely changed and we've seen a lot of that now on parallelization you know we've seen mpi really taking off we've talked a little bit about the different forms of mpi um but the specification has definitely evolved over time but that specification uh has created a lot of compatibility so even though we have many different mpis coding for any of those mpis should be the same no matter
what right you're if you're if you're adhering to the uh to the strict specification um some there's some alternative or optional pieces of the specification but as long as you adhere to the strict specification it should be mpi compatible no matter what now every programming language is going to have its own uh aspect of or interface to mpi so you could have for example uh c bindings for tran bindings for trend 90 bindings and more recently we've seen a lot of interest in higher level programming language like python and python is absolutely taken off in scientific computing uh to the point where uh i actually
think it's probably one of the most dominant programming languages nowadays and i'm kind of curious from from others are you seeing the same thing and what is that like in industry specific areas like life sciences is it similar there too our software stack is almost entirely conda at this point like that there are some there are a few applications in there bow tie sam tools that kind of stuff that that is not in python but for the most part got to be at least 70 or 80 percent of everything is done in python maybe with a mix of r in in that too glenn are
you seeing the same thing in in your area of life sciences yeah things have evolved from from pearl to python python has really done a fantastic job um in uh adding libraries and then and then of course conda um to support uh the researchers it's i think the most popular language um there is a lot of r in there from from folks who come from a statistical side they you know they seem to you know like r a lot but python is is you know this is just my opinion but if you're trying to do a lot of other things than just statistics python is
you know much better you're not you're not bringing up a you know a web server in r so we see python doing a lot um in in the life sciences um the the folks that are using c are the folks in writing the algorithms whether it's uh you know for dna alignment or you know uh drug docking or uh or on the gpus trying to you know get their stuff uh onto the gpus with cuda um those are the folks using the low level languages everybody else is is a little bit higher up at the python level um there seems to be a little bit
of a push to i think i can't remember the end of the project but actually getting some actual typing as opposed to just you know static typing as opposed to just dynamic typing uh with python for some things in uh in research so i'm actually kind of interested to follow that um but yeah i i that's that's what uh the trend has been it's it's funny for me even to consider and think that hpc has changed so much over the years that we're not even focused solely on compiled and and purely efficient languages that we're running across a whole bunch of uh different languages um
i think there's there's even some rust that i've heard of um patrick i think you've mentioned rust as well uh kind of coming into hpc there's there's rust um and also i mean there's still there's still some rubies creeping around in the eda world um and moving moving in there and uh let's see there's but but primarily it's these days it's it's python and and rust and again some some ruby and and of course sadly that pearl still exists um this was awesome you're making greg cry on her no i love the swiss army chainsaw but i also hate it um you know if you've
ever walked into a project where where you've had you know somebody who's they've been in their own pearl world for years and years and years just designing it and designing and designing it and just adding on and then you try to take over that project and figure out what was actually done with this stack of pearl scripts it's um it can be fun and then you rewrite it and fashion like 10 lines and make you hate pearl yes uh dave so we have a new question from dave i've seen so much scripty and numpy usage at scale these days there are a lot of compile
time optimizations to be considered for those libraries alone intel compiler and kl scipy by the way what did i say sorry skippy no sorry i heard skipping maybe i heard it wrong i've just i don't know i don't miss any opportunity to poke at you zane and because there's not many you didn't know whatever i did i see um do you have a question or is it just like to make fun of zane time i'm good with that either way well as i said it doesn't happen often so i have to take every advantage that i can well as on that it happens all the
time because he's good at it too i got my shirt on exactly look down yeah yeah so um yeah there's there's a lot of there's a lot of scientific libraries that are that are also becoming more and more available for python and other higher level programming languages like scipy numpy and and and others i haven't seen and i'm curious if others have seen uh architecture optimize math libraries uh like mkl like intel mkl and whatnot uh are there interfaces for python for these or are we kind of moving away for the for the kind of specific need of high performance and and absolute tuning and
profiling and building the the fastest programs ever in exchange for kind of more general purpose uh hpc applications and and and whatnot i'm kind of curious i know we're kind of at the top but i'm really curious on people's take on that maybe that's where we leave off for our next roundtable it's got to be at least seven or eight years since i built a blaz library to do specific optimizations i think at some point the hardware got fast enough that people certainly at the level that i have to deal with they don't care about getting the last five or ten percent out of anything
right like the next generation of processor will solve that problem for us they focus more on the writing a better algorithm in general rather than trying to get a little bit more compiler optimization out of it when uh when intel came out with their with their uh like intel python and then like mkl and blas and boost and things like that optimize underneath we were really excited to use it but we it just didn't seem to play well with the typical like conda at pip tools uh we so when we try to install it like it just kept blowing up on us and so we
we gave up on it because we thought we we could get a lot of good performance gains because uh everyone was you know using python but just it came down to you know just usability with or just trying to introduce it into already kind of a mature you know python environment that everyone was familiar with hey just try to use these libraries and things like that it just kind of didn't work so uh we we eventually dropped that um but yeah we other than that we haven't built um we haven't built or compiled anything specific with uh for for tuning or a library in it
yeah in a really long time yeah my my most frequently used c flag is m tune generic just because i want it to run everywhere right i don't care if it's fast i just want it to work everywhere i'm going to run it that see that's that's that's what really comes down is that you can make snowflakes that work really really really well one time making something resilient that can run everywhere all the time and reduce your issues is is so is so important to hpc in my opinion because if you're constantly retuning retuning retuning for different architectures and different like that that last two
percent of optimization then how much time are you losing to get that two percent when you know you could just throw more hardware at it for more course throw more memory or you know or throw a cloud at it i mean that's the other aspect is now that you can you know fairly easily scale up it's kind of like for for some of that it's easier to do that than it is to optimize like just down to that little little tiny little bit at the end thank you for that patrick we have one more question before we wrap up uh are there analogs to mkl
for things like the arm ct for architecture so i don't have an answer specifically for the question um i would assume if there's not already that they're going to be coming they're coming soon um but i just wanted to say aydah dave it's been a long time since we chat it's great to see you here and thanks for all the questions dave does anybody else know if there are uh optimized versions of math libraries for um arm dave we're going to probably have to reach out to some people that are running some arm clusters i know sandia has some pretty good-sized arm clusters um and
uh and as well as maybe some arm people and we will hopefully have an answer for that either in our in our in the chat or uh maybe for our next our next webinar i'm taking notes well guys we are a little bit over and we do appreciate your time we know you guys are busy and thank you for joining us again krishna nice to meet you gary john patrick glenn forrest as always craig thank you very much guys go out like subscribe and we will see you next week thank you very much you
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.