Code to Conquer: Best Practices for Achieving Peak Performance in HPC Systems webinar poster

Code to Conquer: Best Practices for Achieving Peak Performance in HPC Systems

Watch Now

Unlock the full potential of your High-Performance Computing (HPC) systems with our upcoming webinar, "Code to Conquer: Best Practices for Achieving Peak Performance." Dive into the world of optimization as we explore key factors such as workload balancing, resource management, and code optimization. Learn proven strategies to elevate your system's performance to its peak, ensuring you get the most out of your HPC infrastructure.

Transcript

[Music] [Applause] [Music] [Applause] [Music] oh [Music] good morning good afternoon and good evening wherever you are thank you for joining at see CQ were focused on powering the next generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and aper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source happy Monday Monday geez a Lise over here I actually had to stop and think about that for a moment like wait where am I what time is it what's happening Thursday

it is the 18th of January 2024 the 18th so I just want you to know that people can actually see you behind the scenes before we come live here I was just notified that someone was watching me rock outq music that's super embarrassing but also funny and awesome I'm pretty sure you've been warned you know that's true that's true then they're always watching right always they know who they are that's awesome they know who they are so speaking of everybody watching yeah we've got like a whole cool crew coming on here very exciting hello my friends Gary so good to see you again it's been

a while Gary how are you Happy New Year than all right so you actually have a really nice long description of of what you do I love that do you want to just kind of like say hi and tell people what what you're doing in the world me me yeah oh well um I'm I'm the science it department head over at Berkeley Laboratory and the portfolio things that I do include managing the institutional high performance Computing and related services for scientific Computing and I'm also the uh HPC manager for institutional HPC down at UC Berkeley awesome very cool yeah glad that you're here uh Brian

you want to say hi and tell us who you are hey everyone good to be back on the webinar my name is Brian fan I'm a Solutions architect here at ciq backgrounds in HPC admin ministration and architecture and uh excited to talk about uh benchmarking and performance today yeah Jonathan man my story is the same as Brian's I'm a Solutions architect here with ciq background in HBC excited to talk about performance that's awesome next thing you just be like yeah yeah what he said what he said ditto Forest what's up man good morning all oh my gosh I just realized up my microphone good morning

View full transcriptHide full transcript

everyone uh for singing whoops um good morning all my name is Forest bur I'm a Susan's architect here along with Jonathan and Brian so you know what these guys are up to see very cool very very cool and zann I mean I don't think you need an [Music] introduction oh dang Allan in the house too all right you got to introduce yourself now on just myself to you guys yeah I think you could write my bio at this point oh you want is that what you would like me to do okay next time Alan I will come prepared and read it for you oh I

we should do it like Stephen coar does with the puppy introductions right you know I make it up we could do that that would be fun too like where we met right and just like make it up where do you know this guy from well we were on the Moon having tea yeah right Call out the wrong University that can be [Laughter] fun you don't mind Alan it's 2024 we have to reintroduce ourselves right I know Happy New Year yours yeah everybody else in my life seems to have Alzheimer's it's like the the state wants me to register my car again you know they forgot

that I registered it last year and then there's the training oh don't we started a training Gary knows what I'm talking about it's already time for awesome Gary's laughing he knows he knows okay well talking about today Allen is awesome and he is head of the HPC department and a good partner of ours at Texas Tech and what is the NSF CAC uh oh well so um yeah the uh National Science Foundation cloud and aonic Computing Center um I had to practice for two years to learn how to say autonomic uh but basically it was a combination of two other centers one focused uh one

we proposed focused on cloud standards and another that was working in the area of autonomic Computing an area that IBM kind of popularized the idea was self-healing stuff uh self-managing so we just usually say cloud and highly automated Computing awesome thank you for that thanks everyone for being here um very glad that you were here so today we're talking about code to conquer best practices for achieving Peak Performance in HPC systems so I have just kind of like a lead question is it always speed because that's what I think of like Peak Performance Allan came off mute real quick yeah uh I'm just whenever I

think about performance I think about reliability okay uh one of the uh key barriers in a lot of especially um you know University scale systems is keeping all the hardware tuned up and functioning um the point where it's not tripping over itself I I actually don't know how uh the really large outfits manage this though I have to confess to having been very impressed when I went through uh Microsoft's Azure um luster setup uh so you know reliability in the sense of um being able to find and eliminate things that you know error rates might cut into your performance in ways that never let you

achieve it I guess the place that a lot of people encountered this for the first time is they buy their brand new cluster and they turn it on and they want to do an hpl run you know or hpcg or you know benchmark it and what that one note is just acting up you know why why don't you just lock it out of the run and you publish that result so you know um uh yes uh you know uh just draw speed is is important but to get to that speed you have to have highly tuned up systems scary yeah I'm I'm just gonna chime

in with Allen and just say it's a proxy for a lot of things but maybe uh I I mean I don't know this is a good analogy but it it's like getting a say a performance sports car or something and it seems fast to you but is it really that fast and then in this case um you you want to get your money's worth because you had purchased you know you spent a lot big investment and if it's not performing then uh then you're not making use of the investment and so in fact if it works for you then you probably could have got away

with a smaller system if it wasn't performing at Peak Performance so uh Peak Performance is just kind of a proxy for a number of things which uh you know some of which Alan Alan enumerated Brian you said benchmarking first what are your thoughts uh so when I think about uh performance uh I mainly think about time to results from like a user perspective and this could be whether you're on the Enterprise side uh doing product development like having a faster time to result it gives you that Competitive Edge to be first to Market and on the research side you know getting your results faster allows

you to put together your paper together and submit it and basically if you're and you know these papers have deadlines and you want your results fast to meet that deadline basically Jonathan you're smirking am I smirking you're smirking I don't think I'm smirking I think you are okay um so all of this is true um Allan and Gary's comments both about kind of systems reliability and actually being able to deliver the performance of the system has available um or or you know in theory has to me is uh founded on this this understanding that an HPC system is not just one thing with one performance

profile but it's a large interconnected complex system uh and that can lead to things like interest uh failure scenarios or or unexpected problems where some part of the system goes wrong and maybe you cordin off that part to to get the performance you want out of the part that's working well uh but it also means that there are many opportunities for kind of very uh subsystem dependent bottlenecks within your system and you know when you design a system you try and make sure that it's well balanced and you don't have like a really coarse example might be you don't have CPUs that are very fast

but not have the memory bandwidth that's available to keep the CPU fed with instructions and data to be processing but that same thing is true throughout the system from uh things as normal as uh the the data system the io system either well maybe network but I think a lot about um storage performance I had a site that I was at one time where we like 10 xed their performance by going from a uh a disk based gpfs system to a flash based gpfs system this was at the very early days of where nvme might be what uh a thing you would even consider doing

and we had some somewhat bespoke um hardware for doing that in this environment you know my like wanting everything to be perfect mind wants to say just rearchitecturing kind of a bottleneck from the system side or from you know if you're on the application side you can try coding around it but there are also things as as mundane as the queuing system can introduce bottlenecks so if you're I think a lot about the difference between uh high performance Computing where you have like one big job that's going to run across a big chunk of the system or the whole system and what in my mind

is high throughput Computing where you have tens hundreds thousands of small single node or even single core processes that maybe run for a very short period of time and in that high throughput Computing workflow especially um your queuing system can introduce significant bottlenecks just in the scheduling and getting processes out to the machine so that there can even be something that runs and this is something we don't generally think of as the performance of the machine but to Brian's Point time to result is really all that the user cares about and they don't care whether uh the system is slow because it took a while

for the job to get out onto the system and run or if it was because it was slow when it was running they just know how long it takes for them to get their result yeah an example of that would be a typical learning curve for people uh running on a on a cluster of uh you know suppose you're launching a large number of jobs do you do them as array jobs do you do them as uh uh as you know how do you count the number of cores uh and optimize that what if you have variable um core requirements at different phases of the

job um so yeah I think this all gets lumped under your scheduling thing but uh but um but I want to go back and and emphasize reliability once more time because I know I said it before but you know um when we invested uh over you know a million and a half dollars on uh improving our power and cooling setup um we just had better up time and that's performance also you guys keep taking all my bullet points here whenever I was thinking about this I started writing down the the different things that I wanted to hit on and I think now you hit on

most of them we have to dive into them Forest before I start diving into that do you have anything to add nope for the most part I would just Echo what everyone has said here um in my experienced users uh when it comes to the performance of the system are mostly concerned with being able to make efficient use of time before as Brian notes those deadlines for papers conference submission stuff like that um being able to have a system that is not only decreasing that time to result but uh is highly available in the meantime um I know I always remember you know the cluster

goes down for and this was very rare but on the rare occasions you know the cluster went down for a second we immediately had people you know banging down the door like hey you know I've got this deadline coming up what's going on um so yeah just like Allan notes you know even investments in cooling stuff like that whatever ends up not only speeding up on you know a core you know individual system faster memory that type of thing um that's definitely something to consider like those things are going to directly affect the speed of the result um but then those int I shouldn't say

intangibles but those other tangible um things that increase just the availability of that speed to everyone is just as equally as important as having a fast system itself so I'm G to read you my list of ones that I put together and one of them I put in there I wouldn't have thought about it but I started thinking what is Allen gonna say or what would he what something that he would go to so I had code optimization which I think we've already hit on iO optimization Jonathan you hit on that one workflow optimization which I think is important for everybody Energy Efficiency was the

one that I was I knew Allan would probably bring up and he did uh and then monitoring and profiling those are the ones that I really thought were important so let's dive into each one so I know we talked a little bit Jonathan about IO optimization I think there's a lot of things to be talked about around that and where data lives can impact that um let's talk more about that if you don't mind yeah so like in in the olden days of my experience which you know isn't isn't that long ago uh but a big part of it was separating out your metadata from

your data storage and this was especially when you have um you know relatively low performant from a an i or from a latency perspective um disc storage you want to save your dis for kind of large bulk well-characterized IO and your file system would help with that but uh maybe you have tens of thousands of files in a directory and that's a lot of load a lot of small iops load on your file system so uh major parallel file systems like gpfs and luster would give you an opportunity to separate that metadata out to something that was better for that latency so uh originally that

was maybe you had 10K SAS hard drives with a a raid one mirror um for that or you know a stripe of raid one mirrors a RAID 10 uh to do that these days almost anyone if not literally anyone would put that on uh some kind of NVM storage just to get that near memory speed uh for doing that kind of metadata operation but you know increasingly it's kind of all going flash we see uh you know one of our partners vast has a really cool story around uh how to characterize with your file system your workload to fit well on different classes of um

of solid state dis or solid state storage I guess and uh so a lot of that is going away but it still is useful to just keep track of from an application standpoint what IO your application is doing and how it's mapping to all of the different processes you have running if everyone is kind of thundering her onto one location or if you're spreading your med data load out you know I I don't have as much experience from the application side I'm very interested to hear what what Gary and Allen's experiences are in this I just see the horror stories of a luster system having

been knocked over and lying on the ground in shambles yeah so I feel like this sorry I thought I was being asked go for it I almost feel like we need a GitHub repository to go with this um this series so we can share code uh we I have a bunch of scripts that um just uh sample IO from each node and aggregated uh up and we've been working with a vendor um to get uh MPI uh versus uh dead IO uh bandwidth separately um that's not quite ready for release yet but uh in general I think we have a lot of tools for CPU

profiling even GPU profiling but not many tools for Io profiling that you know stand the tested time um other than just benchmarking the the the system um we actually just upgraded um the I monitoring software and we bought a new storage system using a vendor product so I'm looking forward to that um slurm Can report IO statistics many schedulers can report iio statistics uh but I found it better to have you know real-time monitors uh that uh that give me you know I can actually ask the cluster you know which account is um drawing the most IO and from which nodes and what are the

highest IO commands U so yeah like I said we should have a a GitHub repository where we can share these things Gary well um yeah the um yeah we've noticed over and I'm sure everybody has but uh over the last five years you know the mix of jobs has gone from like which used to be mostly MPI and some python to like almost mostly Python and then the way python hits the parallel file system with all his files and an environment it just really we just noticed that the high watermark on our parallel file system was like pegged was the performance was pegged all the

time because of all these files from people firing up python jobs so um so one thing that we do is um uh you know have people use thater to stick it into a container and that that helps out tremendously um for for that small IO and then more recently uh we've been looking at um uh We've installed a number of these vast systems and they've been really good at uh like life science workloads which are kind of can be kind of pathologic in their the way they uh tank the performance so um but they but out of the box without any tuning we' we've had

really good success with the vast systems um for using life on life sciences workloads and uh and then I was just thinking the as we're talking about that while Jonathan was talking about like how a scheduler uh can make it performance can make a difference I I do remember different versions of schedulers you know how fast they spawn the jobs can um uh you know that can make a difference if the jobs are like super short for like genomics work glads so then all of the time is in like how fast it actually spawns the jobs so that that makes a difference um but if

we're talking about code opt optimization you know that's a really tough thing because um then it's more like hand toand combat with the users that you're not turning one big knob and speeding everything up you actually have to go out and and work with these users and so we've been focusing on that at at Berkeley lab quite a bit so I have these three consultants and and we you know we put out things like is your python code running slow and then we're just going to just kind of knock these out um and then to entice people to do this because then it benefits everybody

including your investment um you know we we give out allocations here and uh to get people to come in we just said look you come buy your off our office hours for a tuneup on your code and we'll we'll give you 10,000 uh hours on the system for just coming in and you can come in as often as you'd like and you can get that but you know we want to create otherwise if they you ask them to do it on their own they won't so so that's so we're trying to um entice people to do that so that we can work on that end

of it yeah code optimization to me is interesting just because do you have people that come in that really don't know what they're doing at all they've never done this they just assume they can consume everything on the environment and then how do you combat that I'm assuming you can do that with some of the schedulers you can lock them into a very specific set of things that they can consume but does that always happen Brian's laughing Brian is like yeah he's like Brian's laughing I mean you would you would be surprised the people that we talk to that are respons responsible for their HPC

environment at wherever they work and have no idea what to do with it they're like it's great when it works and when it doesn't I don't really know what to do and that's pretty common right so it's probably why Brian's laughing yeah he's like yeah that's what I do I come in and I fix the problem I spin up a little HPC environment with our HPC stack here at ciq press the button and go I mean that's a little simplified but come come off from you Brian tell us your story um just from anecdotally from my experience uh with on Prem systems uh typically uh

I guess partitioning of the system wasn't done well in the beginning and uh basically a lot of the jobs that are running from the various development groups were just extremely in inefficient and uh just from making changes to like your Schuler configuration there was significant Improvement uh with user workflows they were just getting scheduled a lot faster and you know the users were were happy after that uh yes very very happy very happy running users happy yeah so this actually kind of goes into a question um did you want to touch on that real quick Jonathan because the question actually kind of goes into this

I I I I'm respon so you behind the scenes everyone we've got uh we' got a chat here where we're kind of sharing thoughts and and potential things to talk about and I'm noticing Allan here talking about a number of very specialized things you can do to minimize startup time as one aspect of performance um and it it makes me think about there's always like more bespoke more specialized things that you can do if you know your workload you know your users you know that it fits well for them and and they can do that and that's part of the issue here I think when

we when we have conversations like this about HPC and the question that was just on the screen for a moment is is there like a documented process for doing this right and and getting this all correct in the first place and part of the problem is there isn't just one answer and it's an optimization problem against what it is you want to be doing if you have a single problem that works really well on a single like you know multi-processor but single socket or dual socket system now a single memory domain and that's all you need maybe you don't need a cluster maybe you need

one big server uh or maybe you need a bunch of gpus in a big server that kind of thing um if you need a cluster and you know that you have some MPI workloads let's say that you're going to run you know you can get pretty far by just having some uh omnip paath or some infiniband and some relatively new decent servers on a network and then running slurm and just run and stuff on it uh but a lot of this conversation is run by large sites that are optimizing the last bit of performance out of every dollar that they spent because every bit of

performance optimization that they get out of it saves them a force multipliers worth of time for all of their users on this huge system so I think really the summary of what I'm thinking is a lot of this conversation makes it seem like if you if you don't get it right you're just leaving money on the table if you don't get it right right you're wasting your time or your effort but this can be complex but there's a lot of ways for it to go right too not just a lot of ways for it to go wrong and you can get a lot of work

done and and you know people in research have been getting a lot of work done with a stack of desktops in a closet and you know it's it's always a process of iteration and learning and just making it a little bit better all the time you almost never will get it just right from the beginning yeah good point I saw you come off mute Allan go ahead yeah so I um I want to pick up on what John said I think U you know the question is go from Zero to Hero well hero fortunate is a fairly achievable state in HPC I mean it is

an area that generally speaking if if you uh can it all keep the machines busy it's usually pretty productive I mean just as a blanket statement um I think uh you know academic outfits that have well equipped uh HPC and academic research Computing centers uh and I have a whole set of lectures I I give about um talking to Upper Administration about return on investment and usually all all you really need to do is to document what you're doing because it ends up supporting a lot of research it ends up supporting a lot of classroom instruction and if you just write it down often your

conversations with um with administrators are are are easier because uh also when you list the researchers they they're they're good at spotting the uh the vocal faculty members in the list and and thinking oh I don't want to get that person upset so you know you can use the the system to your advantage to some degree in a business setting you know uh HBC I think also can be shown dollars and cents to be very productive so I think it also goes back to what what John and Zan were saying that uh you know just getting it running can be good enough at some point

then you can learn as you go um you know the the the premise of this particular call though is I think is how do you get the maximum performance out it and I think the question centers on that as well uh so I would give some social advice uh you're not the only person in the world doing this you reach out to other people who are doing it uh you know this webinar is maybe one way there there there forums I guess you guys can put a link to your forums in your in your chat there uh and um beyond the ciq uh Viewpoint that

is you know in Academia we have the campus Champions we have an organization called K CCC campus Academy research Computing Consortium um uh in a very general way uh we started about a year ago the HPC docial set of forms you know it's got slack and Discord and Mastadon and uh and uh and projects and it's even got a good first projects uh page if you go to HBC Doo and look at the project says a maybe I I don't know how to get started doing HBC Computing well you can just you know contribute to one of you know a couple hundred uh projects that

are listed there just find one that you so you're not the only person in the world doing this and um reaching out to other part other people uh comparing notes I think is a good piece of uh and and as important or maybe more important in some settings than actually just massaging the hardware thank you for throwing up those links thanks for putting that out Alan you beat me to it I was gonna say HPC do social do you oh Jonathan Maple hey it's also an optimization of the users too and how they choose CH to leverage what is deployed like isv versus recompile and

go or re architecture architect yeah and alongside if if I haven't missed something there um a lot of that goes with training too and I I think a few examples of best practices and and I know like that's the point here Gary was talking for example if you have a python code and it's running slowly these are reasons that it might be slow um here's a possible solution use obtainer to containerize it because it per helps the small file IO demonstrating that to users and showing the benefit I think some of the one of the issues for me again coming not from an applications perspective

is how do we even know when something is slow maybe someone's running their code and they're satisfied with the performance and they don't know if they made a small change that they could get a 2X performance Improvement um have any of you ever had any experience with how how you help a user identify whether their code could be optimized or not th this is a really really good question I think a lot of people are curious H H how do you know Forest you you have an idea I recall that after about three or so days of a job sitting on the queuing system from

a user who I knew was very new on the system I kind of got an odd suspicion that something might be up and sshed into the nodes and found that it's sitting there you know via htop that there's nothing it's not doing anything the CPUs all the nodes are completely idle it's basically just taking up you know three empty nodes um so I got back in touch with them was like hey just you know I canceled this because this isn't doing anything um and we ended up getting back in touch with them got it optimized got it working but just kind of like being aware

of your users and their experience levels and kind of who they are um this in particular was a user who I knew had just come in on like a summer research program and so I was like oh this I about guarantee this person knows very little about HPC um so just kind of you know keeping an eye on those new users um you know looking for those kind of you know smells of something wrong like things that are sitting for days on end but that don't seem to you know be completing there's no new jobs coming up you know things like that it just gave

me an odd feeling so um but that type of thing knowing your users uh knowing who's new um knowing who might need help and kind of connecting that to what's going on in the queue Gary um yeah you know as I mentioned we we have Consultants to help help users and so and we just find a lot of in our case they researchers or scientists and their specialty is not programming and so they're not familiar with oh oh these different versions of python libraries which would be better for them and so we a lot of times we could able to have quick wins by just

say oh you should use this instead and then we just recently did somebody something for somebody and got a 33x speed up so um so yeah I I think a lot of people who are are nude programming because we're reaching you know everybody's using Computing now and uh especially with AI we we do help with people using which AI framework and you know how do may want to do that we get a lot of Consulting questions about that too so that's a that's a big part of what we do now is is the Consulting um and then on the user side uh for user training

I heard somebody mention something about training um what you know a lot of people they they may know how to submit the job but they don't know a lot of the other commands so we have specific training courses where we give like tips and tricks you know how to get the most out of the scheduler and that's helpful too because a lot of times they don't you know they they overestimate the time it takes to run the job when and they if they had just submitted with a shorter time they could have got the job to run so uh that we've gotten good uptake on

on that class too but yeah I I I think that's um it's it's worthwhile people appreciate it people really appreciate it Gary's got sorry Gary's got a bigger outfit than I do but uh but we we rely on a lot of automation to spot problems for uh we have just handful of Staff uh and yes we do the sort of thing that Gary does and I tell them you know within three sentences the researchers are going to expect you to know their entire field and why their code is running slow or why they're stuck on this uh and if you if you don't like that

kind of challenge you should should find another way to be happy but uh do this is the spot for you um but sorry I interrupted you John well I I think you've actually kind of given the next step there and and what I was going to ask about a follow-up you said you have automation that helps you identify people that might uh benefit from that Consulting can you talk more about that Automation and what it's looking for and and how that feeds into your process yeah it's been a work in progress uh at the at the cic the cloud and autom Computing center with industry support for about uh half a decade now um it started with some um some vendor work on uh actually on performance um and I gave a link I don't know if we posted it but it's uh the ID data visualization lab.

github.io hbcc uh this shows some of the visualization and automation work it mins uh data from um the baseboard management controllers at very high speeded using redfish uh Telemetry so we get uh instead of once every few seconds we get you know uh you know up to you know 100 Hertz of of information if we want it um we dial it back if we need to um but it's done by Telemetry so it it's server push uh and it doesn't have much impact on our Network at all and we combine that with scheduler information in various ways and the various um uh you know my

my experience working with visualization scientists has been amazing because they find ways of looking at your data that you would never have thought of and they have all these Sand Key graphs and radar charts and stuff like that and so um it's mostly been a demonstration project of what you can do but for example you we were able to spot um a researcher who whose jobs uh uh were very compute intensive for a while and then it waited for like 20 minutes and then they were very computer defensive and they were big jobs that took a lot of nodes and we found that uh he

was doing something not very clever with the io in between these computational phases and we were able to improve that um so this is automated in the sense that it automatically sorts by user no matter where the user's jobs are across the the node or you can find all the nodes participating in the job or uh and um sometimes they'll they'll come up with a graph and I'll think how on Earth would I would use that but uh uh I I just put it out um we also gave a talk a couple of years ago at supercomputing on this that was very well attended uh

and um the code is available on GitHub as well so um now the gotcha is that uh we do depend on this high-speed monitoring for the baseboard management controller and that generally speaking for most vendors is something you'll have to upgrade your license to get um uh one of the annoying things about dealing with vendors it's true so there's a question from Dave Dave welcome good to see you again I think it's an interesting question too because it also kind of drives the flipping this entire conversation upside down and saying is it different if we're talking about Cloud there's a lot of HPC stuff that

happens in Cloud that maybe it's not viewed as HPC especially around AI today but his question is do you find that more jobs are run inh on in-house Hardware least Hardware Farms or the big cloud providers for each of these use cases Forest I know you're interested in this AI Cloud a lot of the stuff in AI is going on in the cloud just because kind of as that scales up so rapidly that's the easiest way to get access to the underlying compute power that people need for that um actually like you know getting like an 8gpu system and stuff like that of you know

the latest gpus is quite expensive when you look at capex versus Opex of just like renting one online for a little while um so a lot of that stuff is in the cloud um that's always kind of a Byzantine problem because the cloud is always configured you know by default to cost you the most maximum amount of money um and so for a lot of small HBC sites that ends up kind of being a barrier to entry um you know figuring out how to Tinker with all these different options and stuff to make sure that they're uh being able to minimize that cost whereas you

know lot these people are far more familiar with um kind of the financial setup of building one of these big systems like that um so overall these days like I said with AI uh just kind of it's a lot easier to just rent something online than it is to build a dedicated like massive AI system so a lot of that's online um even like p100s v100s some of the you know best gpus of yester year are rapidly aging out of um usefulness uh uh and so yeah in that case it's kind of a lot in the cloud I know we talked about this before Gary

Allen from your perspective are you guys using Cloud I mean you have pretty large on-prem environments but are you are you bursting to the cloud are you leveraging Cloud as the place to run specific workloads what does it look like I should let that Gary answer first my answer is is no but but I want to hear what Gary has to say uh we we are Cloud uses increasing so we actually do make good use of the cloud uh I I think um though it kind of runs not in conjunction with our onr let's just say it runs parallel to the to the cloud because

people don't burst out from on Prem to the cloud they just kind of um you we we don't like having institutional thing set up in a cloud for for people other other than the fact that we may have like an AWS organization so we can manage all the accounts but most of the people are just using the cloud tools to spin up something and and running it that way whereas um it makes it harder because we don't we don't know how they're running their job we don't know you know how it's performing we we don't see other than the cost we just maybe see the

cost um so a lot of our focus is on theond Prem and and because the institution considers Computing essential component of supporting research than than it is subsidized so it is a big cost Advantage for researchers to use the on on PR uh resources because not only get it is it cheaper but they get the support from us to so some of the people I've talked to use the cloud first in some cases because it's where the data lives if it's shared data from someone else and it's cheaper to execute there than it is to bring it in house is that still the case yeah

we we see that yeah because a lot of times it's really fast with especially with our Consultants they say oh yeah yeah of you just use autom ml we just get this going here and and then they'll try it out and say oh yeah that works out great but and then and then maybe maybe and we'll let them go and then maybe nine months of get they'll show up and say well actually I'm using it a lot it's costing me a lot of money and help them move it on yeah so that's been my experience that the butt is that uh I've requested uh a

a cloud research budget for three years running four years running uh never gotten it approved by management because I've done too good a job of convincing them that we deliver better you know Roi on on what we do uh sometimes Your Own Worst Enemy I you know I actually run a whole Cloud Center and we we work a lot with companies helping them use the cloud and uh can't do it at home uh what's the old saying a profit is out without honor in his own country or something but uh that uh I think Gary said it right that uh you know people will go

and experiment um I do want to push back against one thing we've said a couple times that bursting to the cloud is easy uh I've heard I haven't don't have direct experience in this but I've heard from several people who've been trying this that it's not easy for them to burst that you have to make especially if you're trying to use spot instances you have to make a lot of licated reservations you have to get AWS or whoever Google or whatever approvals to scale uh perhaps they've been burned too often by uh customers inadvertently running up big bills and U they don't make it as

easy as they used to but and I was surprised by this but it does seem to be recent feedback that um that scaling is not as easy as you might think uh it is something you have to do in a more planned way um than people might be aware thank you Alan J the rose because I was yeah I mean I just was looking at the time because I was going to ask a follow-up question but I was like I don't know if we want to like keep just diving in I we could probably go forever here so um maybe just just to do a

little Clarity so on Dave's question just as I'm understanding it so he's asking about the difference between in-house Hardware least Hardware farms and then the big cloud providers aw Etc we we know what the major ones are and then he says for each of these cases what is the best source for HPC support or is that outside the scope of best practice [Laughter] discussion Jonathan's laughing there's a lot of answers to that one in my opinion but Jonathan so I don't have any specific experience with least hardware and what I have seen is like least Hardware in your own environment so that's more just a

Cost question and it it's functionally the same as as uh on Prem and even even if it's not even if it's someone else's data center or you have it in a Colo that the actual performance characteristics should be the same it's dedicated hardware for you and you're running on it and there you know we we'd love to talk with you here at ciq of course uh for uh your HPC support in either of those environments frankly any of them uh with out as well uh but a lot of it is the same regardless of where you are especially if you're talking in an HPC context

but there are things that are better to do in different environments but the thing that you're TP that I think anyway that you're optimizing for in each of those scenarios isn't just getting the most performance out of it because that's roughly the same in all of those environments you look at your environment you see what you have you Benchmark your application you can make changes and see if it gets better or worse and uh and iterate to success um but when making this decision of like whether to run on Prem whether to run in the cloud to me that's just another optimization problem one of

the reasons people like to develop in the cloud is because they're optimizing for latency if every time you wanted to do a new project you had to wait for a whole new data center to get built and a cluster to get put into it and budgets to get approved and all of that to happen then your time to solution is crazy uh but you can just put in a credit card number and have a new GPU right now to try it out uh but uh when you are running in that environment ultim like there's there's no free lunch right your your research Computing as as

we often are talking about in an HPC context is is an optimization problem for trying to get the most compute out of the least money usually and uh that usually means you're making tradeoff decisions about not actually needing Enterprise grade services at all times and when you're in the cloud you're usually paying for Enterprise grade service at all times um you know there there are exceptions like spot instances and things like that but even still um the way I like to look at the cloud and I think I've said this on the webinar before for research the value of the cloud isn't that it scales

up because when you scale up you can do it more cheaply yourself but it it scales down and you can get the reliability and the performance of an Enterprise grade solution but just a little tiny piece of it for your experiment right now uh but if you're going to run a lot of it then it makes sense to move it to something on Prem and they're going to be performance characteristics and or sorry infrastructure differences in those two environments and you'll have to prototype and Benchmark and and build your your use case against the environment that you're in but it'll be different from One Cloud

to the next it'll be different from one Cloud's instance type to the next and then it'll be definitely different from one on Prem cluster to the next because there it there there is no standard everyone is different and I'm gonna toss in one thing here that about you know getting readily available from the cloud for for some things let just say to let's just say today it's it's harder to get like a large number of the latest Nvidia G use in a cloud in a block to do like uh training llms for example and then so if you went to ask for that then it's

not something that you could just jump on and use today there' be like a couple of few months wait for that so absolutely and certainly compared to the way it once was uh that is the sensation as as you know AI research in the cloud has spun up and everybody is trying to do it um but compared to if you're coming from nothing if you don't have any compute available to you don't have uh you don't have gpus in your infrastructure you don't have a data center uh the latency of a few months weight in the cloud is still lower than a year or more

of procurement and uh and waiting for budgets to be approved in that way it never takes that long Jonathan come on a year kidding so why we have a little bit of time left this is something I I know that Allan is passionate about and I would like to talk about when we talk about energy efficiency and from a performance perspective you could look at maybe not the fastest return of data or the fastest execution of a job but the most energy efficient use of a cluster to get a result let's talk about energy Allan yeah so uh as I say most schedulers not only

can return iio performance statistics but can be set up to return energy usage and I think it was uh uh uh Prof Professor matsuoka uh reckon that announced at the last supercomputing before this one that they were going to uh start allocating um the fugaku allocations in terms of energy uh consumption rather than uh core hours or anything like that and this would yeah would cause people to uh you know look a little more carefully at the Energy Efficiency of their code uh it's a very multi-level subject um but uh one thing we can probably note is that uh in the pursuit of the AI

hype we seem to have thrown this consideration right out the window and instead we're talking about 100 kilowatt racks and stuff like that to be fair you only really need that kind of U performance um for the the inest route the training part many uh inference or other you know derived uses of AI really not energy consumptive or even CPU or GPU consumptive at all uh so it's really the Bitcoin Miners and the AI trainers that are responsible for our planet slowly melting right um or quickly melting uh so uh I guess one of the one way to frame your questions Z is how can

we achieve this you know pedal to the metal performance that we want uh at better Energy Efficiency and uh you know I think there's some very hopeful things on the on the horizon uh or even already being deployed the arm processors have been setting new records for uh performance per watt and uh risk five uh yeah it's like it ready soon real soon now right trademark uh but uh you know there are U performance U improvements and uh you know Nvidia has invested heavily in an arm technology uh and uh fpga so we may see some improvements but now though I I uh I've actually

taken to on my social media chiding my friends pointing out there is no AI there is no AI it's all uh you know obscure scripting with unexplored failure outcomes or plagiar is plagiarism is a service right so quit quit burning the planet up for to do this stuff and get back to work um uh but I think one way to phrase it is you know what the value you delivering to your organization through each mechanism and can you optimize the energy impact of that in some way I just imagined a like a a recycling ability right with all the heat that is produced being able

to harness that and then use it somehow some way anyway someone will figure that out someone's SM I think the the national renewable energy lab uh enr that's down the street here in Golden actually does that kind of thing I think they run their the their hot water from their cooling system out to to do snow melt and things like that that's awesome I mean just really it's a you know it's amazing all the things that are are are created with HPC resources and I think I'm going to cut us off now guys I love you all like to the mood and back it is

really good to see all of your faces you're wonderful and beautiful and I'm so glad that you are here um at ciq for all of you guys that are watching we are more than happy to talk with you like part of uh Dave Rush's question is you know was talking about um HPC support and that is kind of what we do so you are welcome to come chat with us Jonathan bran myself zanen forest and the rest of the team would love to talk with with you guys we do have our traditional HPC stack and again on the website you can get lots of information

there all open source we got the rocky the Apper the werewolf um and of course you know Alan you talked a little bit about um Automation and we have recently as a company just did a a nose dive into Automation and so we would love to chat with you guys about that as well so definitely reach out to us uh make sure that you like subscribe share this with somebody if you guys have a a another question that we didn't answer go ahead and pop that in the comments below we're always looking at new comments that come up and trying to reach out to you

and answer those as best that we can if you have a an idea for a topic that you want to have covered as well so yay we're so glad that you were here oh and we have a a podcast now too right flops and threads That's it man yeah check out our new podcast as well all right you guys well thank you so much for listening we'll be back here same time same place next week have a great day thanks [Music] guys

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert