Turbocharging Innovation: HPCaaS for Enterprise webinar poster

Turbocharging Innovation: HPCaaS for Enterprise

Watch Now

Join us for a deep dive into the world of High-Performance Computing as a Service (HPCaaS)! Our expert hosts will guide you through the fundamentals of HPCaaS, its real-world applications for businesses, and its significant impact on accelerating innovation and driving enterprise growth. Gain valuable insights from personal experiences, and engage in lively discussions during our interactive Q&A session. Take advantage of this opportunity to uncover the power of HPCaaS and discover how it can revolutionize your business!

Transcript

[Music] [Applause] [Music] he oh [Music] good morning good afternoon and good evening wherever you are thank you for joining at ciq we're focused on powering the Next Generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and Apper support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source all all right we are live you guys this is so exciting yay we're going to get into introductions for all of you amazing people Vanessa it is so nice to almost meet you

we're going to do that in just a second but first of all everyone for watching thank you so much I just want to say that we are on the verge of a giveaway an amazing giveaway and actually I'm a little nervous to do this Dave godlove I think you're the man to do this can you click into there and get a a uh like open up the picture for the image of the backpack it's in our little like shared document so that you can then share your screen like while we're talking about this because we are coming up you guys on almost 1,000 subscribers on

YouTube and that's super super super exciting so I'm looking at it right now officially we are at 95 so once we get to 1,000 we are going to be giving away an amazing prize and why I say amazing is because this is something that you can carry all of your prize possessions in you can take it with you to the mountains you can take it with you to the store you can take it with you wherever you go it is a backpack who who do we got it do we have a picture Dave God love is working on it okay amazing so it is wait

what is what happened to my little slid sheet there okay so it is a rocky Linux carart backpack yeah that's cool man that's cool and you know it's like shout out to all the people on the rocky Linux in the community like the people working behind the scenes doing the documentation making all the updates grabbing the code everyone who's using Rocky Linux across the world which is like thousands and thousands and thousands of companies and people thank you so much for all the amazing work that you are doing it's it's really a cool thing when um something so so used and so necessary in the

View full transcriptHide full transcript

world and like you kind of like wear that with pride and it's cool too I love wearing my Rocky shirt this is a ciq shirt actually but when I wear my Rocky shirt um there's always someone who just kind of gives the nod right as you're like walking around in the world it's like I see you man I see you this is good okay so that is very very exciting thank you again you guys so what we're going to do is once we reach 1,000 subscribers on YouTube so go ahead and sub subscribe share um on the YouTube we will then on that day because

of course we're live every week here same time and same place we're going to announce who the winner is and then of course you know we'll put it in the comments and we'll tag them and we'll reach out to them as um however we need to so that we can find who you are okay wonderful so that is that today you guys today's topic is really cool so we're going to be talking about turbocharging Innovation and high performance Computing as a service for Enterprise now this is going to be a great topic in conversation because there are some differing views on that but first of

all I want to introduce everybody before we get into it Dave godlove will you start please hey everybody I'm Dave godlove uh yeah I'm I'm a well used to be a Solutions architect at ciq I have a different role now um not totally sure what my title is yet but um I'm gonna be training people so that's my new role at ciq and Rose as we're going around like before we you know do too like too much else just want to address the elephant in the room why don't you want to share your screen Rose come on what are you doing there I just feel

nervous that it would like share the wrong thing like it would like be our like internal slack Channel or like because I have two monitors and I'm just like no I'm just kidding yeah hopefully I didn't share the wrong thing yeah no it was perfect it was really good I thought you were going to call out my cool necklace I don't can you see it maybe you can't it's got a little Batman on it and my son made it when he was little and so it's like all kinds like it's got like a weird key and like it's all like off colored and I just

really groovy yeah thank you anyway okay Brian fan who you hey everybody uh Brian fan here I'm a Solutions architect here at ciq uh backgrounds in HBC Administration and architecture and uh it's great to be back on the webinar yeah yeah awesome and forest with a haircut good afternoon everyone my name is Forest Bert I'm a Solutions architect here at CQ I've been with the company for just about two and a half years now and uh my background is in HBC Administration where I kind of came out of the uh academic and National Lab real uh good to be on the webinar and good to

see you as well Vanessa I uh caught your talk at FM and uh enjoyed it quite a bit so good to have you on and uh good to meet you okay Vanessa I mean f first of all the beauty of what is happening in your world right now you're like moving your body you've got all these Sparkles behind you I just love it we have actually never met so please unmute and tell us all about you sure hi I'm Vanessa thet I'm a computer scientist at Lawrence Livermore National Laboratory I probably need to say that whatever comes out of my mouth today is not the

views of my employer that's an important note I notice that everyone here is from control like you except for me uh I started out my work in Singularity so container Technologies and my world has blossomed into much more since then I'm I I would say now I work mostly in the converged Computing space I I develop a lot for kubernetes operators um very involved in standards groups and I'm really oriented toward taking kind of the most common you know Community well-known projects such as kubernetes and making them fit for the HPC use case so that we don't have to pay for these special services so

I have a feeling I'm going to be the one with the other opinions on the call but I think we'll have some good discussion today and I'm looking forward to I like it so I shall put you on the spot and only because I love you and also I have already heard many times these other amazing people's kind of perspective on it um but will you just give your you know and as deep as you want to go or just a sentence or two is fine but like your perspective on what HPC is so high performance Computing uh traditionally you'd only find it at sort

of large centers and National Labs it is a lot of tightly compact machines that are optimized for performance so our workloads are for example physics simulations things that you really need scale things that you need um networks that are harder to come by in the cloud so you need something like infin and and you may not even be able to find them on some clouds um and it's an interesting thing to look at the space of HPC as a service because you can go back as far as like to 2018 when Cloud started presenting products that they would call HPC so for example uh Google

had like they had a batch day not a batch day an HPC day um Amazon started releasing parallel cluster and for the most part it was trying to take these HPC workload managers and put them on the cloud um and that still tends to be the case today and I think I think there's a couple of audiences here in terms of uh I guess You' call them customers so there are the really big centers the National Labs um that are doing you know like simulations that W that are not ever going to go onto a cloud right because for many for many reasons there's large

academic centers that also have their own HPC they are probably also not the customer that you'd consider for HPC as a service I think what's happening now is there sort of a middle ground so maybe smaller centers that don't have the the funding to either get their own to buy their own stuff and pay for it on premises um or kind of smaller companies that are taken into the idea of HPC um and they they want to run something that's more more scoped to AIML I think that's the trendy thing today and one thing that would be really fun to talk about is how kind

of the services landscape is really changing because of the needs of these customers so so for example kubernetes originally was something that was created for running things that need to persist so like I'm going to deploy my website services and oh the container the Pod died like we'll just recreate it no problem but um with with AI and ml workloads that came in and that being a very strong customer base all of a sudden there was a need for batch so a need for something to be able to start running and finish running and so that's when the batch working group started a couple years

ago and it was interesting because that was really the first of intersection between like oh something that HPC has been doing a really long time and now something that the kubernetes community is interested in and Forest if you saw my f them talk I kind of alluded to that a little bit um and so that's that's really powerful because it's this opportunity for collaboration but now that we have this shared incentive we have to figure out like how we can best work together because there's there's some things that are kind of moving in different directions so for example a lot of these AI models they

don't need the highest of precision and so you look at sort of vendors that are trying to have trying to please this customer base that is are running these models and they're designing chips toward that of course that kind of chip is not going to be good for like a climate simulation actually it's GNA be absolutely terrible and so we need to look toward the future and say what is the future going to look like for our science because Cloud does when I say cloud I'm I'm generally saying like places that are offering services for compute and that's opposed to on premises where you buy

your own resources if the future of cloud is going to be a huge predominant player in the space which it is going to be it already is then you can again watch that fom talk to get the cagr I think it was like 20% and you know HBC was a lot smaller um we have to figure out where our Niche is in that space and how we can be a part of the story successfully and not for example find ourselves in a future we have the needs to do science but we don't have the resources to do science and so my perspective here is coming

at it from like advocating for this the advocating for the scientist and figuring out what the technology space needs to look like for that as opposed to like you know trying to make money uh or something like that that was a very long-winded answer and I have a lot more thoughts but I think I'll I'll pause right there so it can be more a discussion can I chime in a little bit of course yeah um yeah so a lot of what you just said resonated with me Vanessa um it's good to see you by the way um yeah um so uh yeah I mean I

think that a lot of HPC centers um have been like really avoiding um you know trying to think about how workloads are going to run in the cloud um you know I I think that um you know for a for a long time uh users would go to HPC administrators and be like you know I want to uh run some of my workload in the cloud to um you know kind of skip the queue and be able to or a lot of times also administrators would um would uh you know go to like like when I say administrators I mean like directors like people who

are like up above kind of the actual people who are working on the computers and would say hey the direction of this needs to be that we're going to expand some of our workload into the cloud and we we're going to task you with figuring out how to make that happen and I think for a long time you know kind of the standard response to that was to go try to figure out what it would cost and that's that's hard it's hard to figure out there's so many variables I've been on I've been on the side of trying to do that and there's so many

different variables into trying to figure out like how like what a what a cluster in the cloud would look like and what that would end up costing that you end up with like these huge kind of error bars on your estimate but usually what what it used to look like is that somebody would go and say it's going to cost a lot more money than we currently spend on on Prem and then kind of Kick the Can down the road some of of the feedback I got the last SC was I was talking to people and say and they were kind of talking about the

pressure that they were getting um to move workloads up into the cloud and to figure out how to make that work and I was like what about the cost issue is you know and and the feedback that I was getting was that the cost issue doesn't matter anymore it actually doesn't matter if it's going to be more expensive to run this in the cloud or not it's a mandate we're gonna have to do it um so that that that future is here it's you know it's now uh you know HPC centers need to become become comfortable with the cloud but um another thing that you

said that resonated with me is that you know a lot of times even people who are kind of in these spaces like people who who are in the like the I'll call it the cloud native kind of community on the one side and people who are in the HBC community on the other side um even those people who are kind of near to it tend to look at uh you know both different sides of the equation as just a big pile of computers and kind of assume that they're similar and there's a lot of ways in which they're not similar and it there's a lot

of ways in which it becomes kind of kind of hard to think about how to move a workload from on Prim out to the cloud I'll give one concrete example which once you think about it it's really really obvious but I think a lot of people don't think about it and so it it it's missed a lot um what does a batch scheduler do for you uh what's the purpose of a batch scheduler the purpose of a batch scheduler is that you are in an environment where you have a finite number of resources on your cluster and you have a large number of jobs and

the the task of the batch scheduler is to optimize um the basically how how the usage of that finite resource like how do I fit all these jobs on here in the most performant way to make sure that the resources is used optimally and it doesn't really care about individual jobs right like if you have to wait if your job has to wait in order to to make the the overall system more efficient well then you know tough you got to wait um in the cloud the situation is exactly the opposite uh you have one job and you've got for all intents and purposes unless

you want a a100 uh infinite resources right and so now you just you're spinning up your own personal cluster that is your own bespoke thing and you know Vanessa said earlier you know going back to what she was talking about it used to be that um and actually still is to some extent the cloud providers would give you tools to create your own like slurm cluster or whatever in the cloud and it's like why do I have slurm I mean maybe it makes it makes things easier because I already have slurm scripts and now I can just use I can reuse the same stuff and

you know just um you know send it to my new bbo cluster uh using the same scripts but at the same time it's like slurm is solving a problem that you no longer have right you're not trying to fit a whole bunch of uh jobs into some finite resource you're trying to like optimize the resource for your job it's like reversed now and that's such a basic thing but you know that I think a lot of people miss it but it's like a big fundamental difference you know and then there's the whole I mean I'm not going to talk I'm gonna cut myself off here

pretty soon too but you could go into the whole Services versus applications kind thing that you know um Vanessa touched on a little bit as well and how that it's just to totally different mindset totally different sets of tools that you need to accomplish those things like a lot of people once again they know one or the other and they they assume that either HPC or the cloud is the same is the thing that they already know but once you start digging into it you see that it's like there's a lot of really low-level fundamental differences there oh I have a I have so much

to say to respond okay so um a quick answer to your question why slurm in the cloud I I think the thing is that people to transition to new things need that familiarity in the same way when we first designed Singularity one thing I really pushed for I said I there need to be things that feel like Docker people are using dockor they're com they're comfortable with Docker that that's the way we need to go so that that's the main reason I think um but you're right I totally agree about that now I think one really interesting thing that you said that I want to

point up is this idea about Cloud resources being infinite so I think there's two Illusions here I think the first is that one there's this illusion that running your stuff on premises is free I log into my cluster and I run my 10 gazillion jobs and it doesn't cost me anything it's free but then there's also the illusion that cloud has infinite resources which if you've used Cloud enough you know is absolutely not true it's not that the resources are infer it's just that you are not told what is available and you have to sort of Discover it oh I'm I wanted to get 64

instances and I just waited 30 minutes and I only got 50 and I had to pay for those 50 and all of a sudden I can't run run my MPI workload because I actually needed all of them um so that's that's that's another thing and cost is an interesting conversation because something I hear from from researchers or actually just a lot of people is that there's no price transparency actually cloud has really good price transparency if you've used the AWS cost apis they are excellent now now there's a lot of data so what I do often is I'll like have strategies for caching it and

putting in an oci registry for use later that's just for one's own um I guess for for another Cloud though they give you like the unit of cost down to a specific type of resource and that's challenging because then you have to like assemble them like Legos like you have to take pieces of memory pieces of CPU and then actually put them into instance but everything is there that you need to make good decisions on the other hand on the side of HPC we have no transparency about utilization about costs if we are to be able to do these actual comparisons between like is it

really economically sound to do my workflow on HPC versus Cloud it's not the cost for the cloud that we need it's the cost for HPC that we need to be able to get because we don't we don't have that data so we're not empowered to do it and that's sort of like that's sort of like our fault which is really interesting um and in the beginning you were talking um about uh what were you saying something about how how researchers or administrators are always going and like asking for the next thing and that's really interesting because that also took me back to Containers right like

it feels like HPC is always lagging behind just a little bit like look the rest of the world is using container Technologies and can we as well and I this is this is why I think it's so important for us to be a part of a conversation for a lot of these standards groups as they're going on because what tends to happen is like a spec is developed with sort of the cloud you know use cases in mind we don't show up at the table we don't express our needs our needs don't get represented and then after the fact we develop our own set of

things that oh you know containers need root well we're not going to run that on a multi-tenancy system let's go back and now make five new container Technologies and redo all that work and we could have just been a part of the conversation in the open containers initiative from the beginning and not have to like not have to have dealt with that and I'm glad that you pointed out that um it's not just a pile of computers so they're they're obviously very different and part of the the focus of the work that my group is doing and converge Computing is really understanding that so understanding

when we take our workloads and we move them to Cloud how do they how well do they do in terms of performance in terms of networking so some clouds have really good networks so Aur has does have infinite band quality networks Amazon has the elastic fiber adapter that's very good Google's networks so they they have good networking if you use some of their more expensive stuff uh they also release a standard called falcon which I was really excited about TM was just a standard and I was like is the standard gonna like turn into a thing because I'd like to use the thing um and

so this is an area that I consider to be research that we have to do to figure out how to best move between the spaces and then how to best create this converge space of Technologies to get the best of both worlds for what we need for our workloads and the other really cool Point um that goes back to my fom talk is that it's not always just about like we have to move all our stuff to the class guess what we can run kubernetes on premises too but we do have to do research about how to best do that too because right now there's

there's some problems yeah okay so I'm just gonna cut in here one second Vanessa I knew I liked you from the very beginning so this fom talk that you guys have alluded to a couple of times is that maybe something that we can grab um for people yeah absolutely I can I can send a link it's called the the kubernetes and HPC the bare metal Bros it was fun that's awesome that that the way that you just said that is going to be in my head forever I appreciate that and then you actually also mentioned cloud and I just wanted to pop in here because

I kind of forgot to say that one of our very own Colin vanders Smith is actually right now speaking at our um there was a cloud optimization Summit happening uh at Google and so it's one of our partners and so he's there speaking so I just want to give a little shout out to Colin thank you for doing good work don't knock him dead we don't want anyone um cool so I I actually so Brian I I just wanted to like pop in here because I see that like a lot of like the conversation that they are having here you're kind of like doing this

like head nod thing so I just wanted to ask what kind of your thoughts around HPC as a service are as well uh so I do agree with uh the sentiment that running HPC in the cloud can get expensive and actually getting the workflow to run in the cloud May be easy but for it to be actually performant and cost effective that's pretty challenging so like taking the most simple case like you deploy some instances run an NFS server on them you can run an MPI job on that sure that's cool that might cost you a lot because you're running on demand so now you

want to get some more performance maybe I use something like awf FSX for example and that might give me a little bit better IO performance for example but at the same time my compute is still using on demand so how do I get the cost even lower now uh you know you might start doing things like May looking into the spot market for example and then by placing a lower bid on some maybe older instances if you don't need the results immediately you can further drive your cost down and maybe in business context that might be worth it for a specific business but uh yeah

and then I guess also from a business perspective Ive like if you're going with traditional HPC on Prem type infrastructure that's uh mostly capital expenditure so that's just an asset that you'd probably have on your balance sheet and then if you switch to something like running your workflows in the cloud it then becomes an operational expense and uh yeah and then you know you could do accounting type things with that and for your business to you know potentially make your make the value of your business a lot more yeah so Brian that's a really good point about spot so I I had this really dumb

idea I was just actually curious so HP uh sorry AWS has these HPC family of instances that are explicitly for HBC so like the 7g and the 6A so they they have uh they're single threaded um they're they're very large size and the price is actually very good so I was like okay I'm gonna try to I made a linear programming algorithm that would try to select a pop perie I called it the pop perie algorithm of spot instances and my goal was to beat the price of the HPC instances and I only had two uh groups of spot that actually beat it and it

was so small a difference that it wouldn't have been worth switching over to spot so I do want to say for AWS those instances the HPC Ines are very good price uh to be on demand instances the one um issue with HPC with not using HPC instanes on AWS for something for HPC is that a lot of them by default are um not are not single threaded they they have two threads per core and unlike other clouds that like will give you like a flag like let me change the threading AWS doesn't so you actually have to hot plug the instance as it's coming up

so if you have like something Dynamic where you're bringing up spot and they're going down um and you can't get it in at boot you have to have like a a script that basically like injects another script that anyway it's a mess so if AWS is listening please give us the the flag to change the threading that's that's what I want for Christmas this year oh and uh just one more point to add to this using the spot Market it's only it's mostly beneficial if your workflows are restartable and have checkpoint restarts as well uh so you know an example would be if I'm running

some large simulation I'm like 10,000 iterations in my stuff gets spot killed you know you don't want to start from zero and you you know you probably want to start at the last checkpoint so you actually can your simulation will actually finish at in a reasonable amount of time yeah MPI is like not it's really hard and like 90% of exascale workflows use MPI it's very common in HPC so we can't we can't just like pretend MPI doesn't exist it's the big player in this Cas yeah and I think that this discussion also highlights something else I mean so Vanessa I mean like I take

your point that one of the reasons why um estimating the cost is such a problem for for HPC folks is because we don't necessarily have the other side of the equation nailed down all the time as to you know what does the HPC system actually cost and but that's not all of it I don't think the the other part of it is that there are a lot there are a ton of different permutations and ways in which you could you know use any given Cloud to kind of like uh simulate an HPC system on Prem and those I mean there's there's like you know there's

like I don't know 100 different variables and then you take all the permutations of that and figure out like how many different ways that you could do it and those all cost different right so you have to like and they all have different performance um implications and so when you're doing these types of um analyses you have to like start to you know make assumptions and take some things as for granted and you like you know it's it's really really hard to kind of sit down and figure out what's the optimal way given all the different ways that I could do this to to try

to you know I know that like you just said that you've actually tried to um you know come up with some algorithms and try to figure out how to do that so obviously you understand what a difficult problem that is you've worked on it um but yeah so that that kind of gets in your way a little bit when you're trying to figure out what the cost is going to be because it's like well I mean there's a zillion different ways I could do this so what cost are we looking at here all right Forest what's going on in that mind of yours around HPC

as a service I just had a couple of different uh points to agree on and kind to discuss that everyone has said so far um first off I think that yeah the cost optimization of HPC is a huge problem because in general monitoring has kind of been an issue in HPC for a while I remember you know using like XD mod and stuff like that in environments back in the day and that worked but when you go look at like what Enterprise has with their standards around grafana and stuff like that it's so many leaps and bounds ahead of where HPC has kind of been

at with monitoring that um it's easy to kind of see how uh it's difficult to keep track of HPC and the cost basis there um because the tools are just you know grafana is one of those Cloud tools that's never quite uh you know permutated back into HPC fully um we kind of touched on how you know the usefulness of slurm in the cloud and stuff like that and how um you know we look at these Legacy models uh as kind of maybe what people want um that uh Vanessa you said people tend to expect you know slurm kind of because that's what they're familiar

with um I couldn't agree more that's one thing that I've always kind of run into on the side of um you know kind of looking at this issue through how I've looked at it uh you know the first question I always have got is well how does this integrate with slurm and you know when I ask well why do we need slurm it's well all of my researchers have all this stuff on slurm and they're not going to want to switch it over and so countless times I've kind of had the same back and forth and uh so I couldn't agree more that um you

know like to Dave's Point why do we actually need scheduling in the cloud um well like you said Vanessa ultimately vcpu limits it's not that bantine to understand those um you know once you figure out where that page on AWS is you know it's pretty easy to you know uh um you can kind of start to get some idea there of like said am I going to be able to get 58 instances am I going to be able to get the full 64 where's that limit going to actually land um so while there there is you know definitely some scheduling that you still have to

do there to maintain around those limits um I definitely like I said completely agree that ultimately it comes back in a lot of cases to slurm being um and that type of thing being familiar I I've had the same thing with containers and kind of trying to convince people that containers of the future um a lot of people have been like openly resist IST when we've gone and discussed with them you know well here's how we can take your code and put it into a container and the discussion is really kind of stall that well I don't really feel that containers are proven technology you

know that's it's been very difficult at times to convince people to actually make that leap into containers and so from just a lot of different sites it's kind of very uh it's a widely variable it comes down to uh the level of containerization that you find sites are very into it they've got lots deployed on it they're using Singularity they're using obtainer they've got you know all that type of stuff um some places like I said are very resistant to even the idea of moving their code into something like that so I think ultimately familiarity and kind of a lot of these maybe dogmas that

we have in HBC is kind of what holds it back from um you're really kind of embracing some of this technology um this is kind of uh ties into why was so fascinated with the talk that you gave because I like that uh it's kind of the whole usern netes approach you discussed is a completely different view to the problem um that you know ultimately in the end brings an entirely different capacity with you know the user space being available in kubernetes pods and that type of thing um so it was very interesting to see kind of some of these slurm problems and that type

of stuff uh Sid stepped maybe a little bit uh by some of that work and so I'd be interested if you've um you know kind of those questions have come up and if people have you know kind of where that's landed if people are intrigued by how slurm interacts with your guys' approach over in the batch Sig or what kind of your answer to those type of queries has been so I haven't presented uh that particular component yet in the batch working group um I want to respond to a couple points you made first I think it's really interesting that so many years have passed

and we still have this culture where people are like containers like I could think back to 2016 like being in a room and talking talking to people out containers and getting the sense of them being like Oh is she talking about containers again like we're we're like we're still in kind of that cultural thing that's really interesting the second note is that kubernetes for developers is like is like this toy store it's like I mean you basically can put things together like Legos and I don't think people necessarily even have to choose what cluster they want so you don't have to choose slurm or PBS

or Torque or hyper Q or HT Condor because there's an operator for all of them and if you want to run it you spin up your own you know mini mini cluster mini HPC cluster whatever flavor it may be in kubernetes um so that said I do think scheduling is still really important and especially for HPC because a lot of these workloads really depend on the topology so it's not just for sort of low latency networks but like literally like the physical distance from like some GPU and like a PCI bus like that's important and right now that level of information is not represented in

something like the CP scale and if you've developed for like a custom schedular plugin the design of it is really kind of well it's flexible to have different kinds of extensions but at the end of the day you have an active cue of like of PODS that are going to go in that can be sorted in some way and then you kind of give them to the kubet and the ku's like oh I'll honor that or maybe I won't I'll change my mind and what's interesting in the kubernetes development space is that we we're starting to see that there's sort of this acknowledge um issues

with the schedule especially because even people in the M the AIML space are wanting more kind of intelligence with respect to scheduling so you're actually seeing tools like Q so that's with a K because everything in kubernetes with a K right where um with different web hooks that can actually intercept for example jobs before they get actually assigned to you're given to the scheduler you're actually seeing like scheduling happen by what comes down to be a controller or an operator um and so I think in the next year the next couple of years there's going to be a lot of uh growth and innovation in

the space with respect to scheduling and it's it's a really interesting space to work in as a developer um if you've never looked at making a custom scheduler plug-in for example take a look at it if you've never looked at the architecture of Q also look at it it's it's complicated to figure out the first time um I actually wrote up a document for how it works if you're interested but um I would I would definitely look at that because I think that's going to be a really important component for HPC moving forward using specifically kubernetes or even usern netes is getting that just right

and remind me again was q that component that you had the graph uh up that showed kind of like the scheduling misses on um so that was just the default scheduler and then spin up kubernetes and like submit jobs to it and like YOLO like see what happened and then the next one where it was all ordered that was q no so that was fluence that was actually using taking our flux framework so flux framework is a group of components one of them is a scheduler the scheduler has go bindings the go bindings can be plugged into a custom scheduler plugin like I described and

all our C and then so we're basically using flux to to schedule pods to the cluster and because flux has more intelligence about like okay these these pods should be together I'm going to put them close together and in a group that means that they all got scheduled at the same time they started together they finished together what happened with the default scheduler is like four four out of six would like go in and then the other two would be like oh I can't actually join yet I need to wait for these other things and then the job that maybe would take like 90 seconds

would take like twice as long because it was kind of a pathological pattern how do you spell that Vanessa the Q is it k u e u e u e like basically just change it to a k that's that's like the kubernetes way you just take a word and you change the the whatever it starts with to a k and yeah so it sounds like a bit of what you guys are talking about is this this relationship between um kind of you know cloud workloads and HPC workloads and how they're very similar in a lot of ways and how different companies organizations research universities scientists

can begin to utilize um and we've seen that a lot more that people are having both right like they're not just purely hey we have this you know on Prem clusters and we're doing everything that we need to do here um we are also utilizing resources in the cloud and where is there like a a a a sticking point because we're talking about HPC as a service but does that does the service part include an ability to utilize Cloud resources when necessary okay so I can I can respond to the question um the question is what is an HPC workflow and if you think about

that HPC is really bad at creating reproducible workflow so if you go back to like the early 2010s the first communities that really grabbed onto uh workflow tools were biosciences snake make nextflow Cromwell they were the ones that were like there's these new things called VMS in the cloud and I also have HPC clusters let me make a tool that's going to allow me to Define my steps and to run them and then when containers came around it was very natural to add the container executor backend like I literally remember in like 2016 adding you know container support for Singularity to like snake make for

example um and then eventually next flow actually I don't think that was me that did that I did Cromwell I think um and so what we have now because of that early work is you find Registries of workflows you find entire like you know bio containers is a whole has like over 8,000 9,000 containers now for just biosciences they are so far along with their reproducible workflows because of that so then you come to the HPC community and you're like okay guys like what do we have what's in what's what's where where are our workflows and here's what you find you find um you find

like Benchmark sites so like Coral 2 benchmarks um I shouldn't call it specific ones but that will take you to Old FTP links where you download a zip with a PDF that tells you how to compile it and it's incomplete and it's really hard to do um you also find GitHub repos that haven't been worked on for eight years with you know without any documentation so the the TDR is that we never merged on a solution for workflows and there's many reasons for that you know a lot of these are kind of top secret a lot of times it has to do with the incentive

structure so scientists just want to do the work get it done publish it move on to the next thing they don't have the time or the incentive to really Harden it into a reproduced ible thing um so then to kind of step back so now we're we're having this transition where we need to go into the cloud we have to I think as the software Engineers that are working on this problem on the infrastructure we have to to some degree figure out what we want that future to look like and so a potential future I'm not saying this will be the future is to say

you know there could be a future where we're running kubernetes on premises and on the cloud that means that that that just having something in yaml we all love yaml right having something in yaml and in containers becomes this vehicle to move seamlessly between the environments and so that's a potential way to do it let's try to investigate what it means to do that um I I also would like to see State driven workflow tools if you look at so um the the creator of nexla has this great repository with like all these workflow tools most of them are dag oriented and I'm wondering where

State machines fit into that but anyway so that the answer is that there isn't a an HPC workflow we have to figure that out I think it's going to involve containers I think it's probably going to be agnostic to like a specific tool so for example um snake make can can deploy abstractions to kubernetes so can next flow people that like using snake make should be able to use snake make next flow next flow um we shouldn't we shouldn't like subscribe them to only one tool if that makes sense I like it see Dave you coming coming off a coming off a mute what you

got yeah so yeah um um yeah a lot a lot there so back to your original question Rose I think that you were asking you were you kind of you know postulated that we've got cloud and on Prem both at a lot of different institutions and so like what's the sticking point I heard you say like and I and I I you know um to kind of uh bounce off of some of the things that Vanessa said as well yeah there's these workflow languages is like Cromwell and snake make and so on and so forth Whittle and you know what well in any case um

the these workflow tools uh like next flow and so on um so in a previous life I was a staff scientist at the National Institutes of Health and you know uh one kind of request that we would get from users a lot is like got this workflow that is supposed to run really great on AWS or you know gcp or something through you know nextflow or Cromwell or something can we get it running and you know it should be easy can we get it running um in the HPC and like that's the goal and in theory you know that that should work but in practice

theory and practice don't line up and so in invariably you know some some poor staff scientist would have to sit down and you know try to Cobble something together and figure out all the bumps and try to you know get this thing and a lot of times it wouldn't it wouldn't happen and so at the NIH probably the same at a lot of different institutions yeah there's a lot of cloud usage there's a lot of cloud resources and there's also an on-prem um HPC actually a few on- Prem HPC systems um and the sticking point back to your your um your question Rose is that

the user interface for the two are different and and because of that it's different groups of people that are using those two different resources um and there's not there's not a whole lot of overlap between them so yeah it really I think that really what needs to happen is we need to develop a user interface that doesn't like in theory work both on Prem and in the cloud but actually in practice also works um on Prem and in the cloud and I think you know Vanessa kind of was talking about that you were talking about that a little bit you know and using for example

the idea of kubernetes um you know that's that's so another another kind of thing is or another you know um idea to consider is that you know kubernetes is an incredibly powerful and flexible tool with a lot of different aspects to it and you know um that makes it hard to learn and so you know I I think that if we're going to do this right what we need to kind of do is we need to figure out um we need to figure out a user interface that's going to work you know and and solve the problems that need to solve and and do what

it needs to do both in the cloud and on Prem and it needs to be simple it needs to be something that a scientist who has a day job as a scientist it doesn't also have to have a night job as a computer scientist and figure out how to use you know and so that's that's basically exactly what we have been working on and are trying to develop and are trying to socialize and uh with with the fuzzball tool that you've heard us talk about a lot on these webinars and that you can see information about on the ciq website and so on so that's

that's kind of our um our uh attempt to solve this problem by having a a user interface which is you know the goal ultimately is to be um a generic user interface that it doesn't really matter what the resources are that you're running on um and also to be you know simple and um you know to allow users to just kind of generate you know Define the AML or if they're not into that use a guey and create an AML from the guey and uh you know be able to run um complicated uh powerful you know kind of deep analysis without having to you know

go take a class and and how to use this other tool you know a yamel class I don't know Dave I I love yam no I'm just kidding and I'll add two words to that that I think are important so the glue I think that is going to hold these things together because at the end of the day there's going to be a lot of environments there's going to be HPC there's going to be cloud with or without kubernetes there's going to be just apis you need something that smells like the term that's used now is multicluster some kind of high higher level orchestrator that's

going to accept the work that's going to have a nice user interface hopefully to be able to do that and there's going to have to be a scheduler in there somewhere too or more than one scheduler even that knows how to dispatch things to the correct place um I would say so I'm my my bias is toward open source so the extent to which you can make pieces of fuzzball Open Source and really bring in the community to work on it together and allow people to kind allow it to kind of grow organically you know kubernetes is successful because of many things but also because

of of the fact that it's open source you know with I don't know what 3,000 active contributors at any given time I think that's a really powerful thing um so I'll just I'll throw in my my opinion about that there I like it we can certainly appreciate that I mean you know we we provide support on on many open source projects and are are deep in there and it is a weird line right right I mean because you say like oh it's not about money but I mean you know it's nice to eat and pay your rent and you I I like eating too well

you know I mean I'm not gonna try to I won't say anything here that is gonna be like uh you know in the future people are gonna be like but you said but you know like look at pot man for instance um podman is open source and it's awesome and a lot of people contribute to it it wasn't always open source right it started off as a project in red hat and it kind of incubated and got to a certain point um you know where it was mature enough that the developers decided to open source it at some point I think that was March or

April 2018 I remember yeah that's awesome all right you guys so we're kind of coming up to an end but I want to hear kind of like a wrap up each of you of like your final final thoughts on this topic and I I'm going to start with Brian actually okay uh I guess key takeaways uh if you want to run your stuff in the cloud make sure you bench mark your stuff make sure it's cost effective and make sense for your business and uh go from there yes good point Thank You Forest um it's great to see kind of different groups of people working

on the problem and it's great to hear different perspectives and to see like I said novel techn ology and stuff being developed as kind of uh the result of everyone working on it um and it was great to have a discussion about it today so like I said great points about Cloud uh and kind of considerations of the old versus the new and uh it'll continue to be a really really interesting space to follow in all aspects so Vanessa thank you for joining us and everyone else thank you for being here yeah thanks for us Dave yeah I think I just sort of reiterate that

I think that the the biggest problem the biggest sticking point that that needs to be overcome is um you know we we really need not just the promise of workflows running um in multiple different environments multiple different clusters in Cloud but we need that actually to be a reality and uh we need that user interface to be easy and accessible to uh you know to scientists that that uh you know need to be able to use this stuff so I think that's really where we need to be focusing our energies yeah awesome thanks and Vanessa take us out girl yeah so we've been talking a

lot about technology I think I'll close with something about culture it's it's been very common that between you know when you think of cloud and HPC there's there's a lot of an adversarial mindset so people from HPC are like oh cloud is expensive and people from cloud are like oh you old dinosaurs what are you doing and I think moving forward to get the best future for both of us is really going to require throwing out that mindset and having one that is collaborative fig figuring out how we can work together being open to new things and maybe if you're a developer if you don't

have bandwidth to directly like show up to a meeting that for the other the other community that you've never been to before you can figure out how to make your project inviting for the other community so for example making a binding in a language that someone could use so just moving forward whether you're a developer or researcher think about your mindset think about your biases and try to figure out how your work can be more collaborative to kind of bring these two communities together I think that's going to be the most successful future for us awesome that was great thank you for uh taking us

out there and you guys I mean that is quite frankly exactly kind of what we're working on here so I'm glad that we had like an opportunity just to kind of you know touch on fuzzball real quick I know it's kind of still behind the scenes being worked on right now but this is definitely the idea that you're talking about Vanessa is really bringing these these two worlds together in a vision of the future where we all have you know access and ability to do all the great work um Tech technologically as we as we want to right and the sky's really the limit there

as we bring these these two resources together um so again you know thank you for guys for watching thanks for Nessa for being here and the guys from ciq if any of you out there watching are um interested in any of the projects that we support here at ciq whether it's Apper or werewolf or um a sender uh or even fuzzball we were still happy talk about that even though we're not like ready to put it out into the world completely go to C iq.com you know you can get in touch with any of us there find out more information um we also have a

podcast uh flops and threads so that is also on our website where you can grab that and get the information there so you can keep in touch with us make sure you like And subscribe and find us on Twitter and uh LinkedIn and thank you so much again Vanessa I have you know like now that we're actually kind of talking I have heard your name in the community so I reach out to Dave I'm like wait Dave is this the Vanessa he's like yes it is oh well she's over there for me so absolute pleasure you are a Divine human being thanks for being here

and all the incredible work that you've been doing with Singularity and and other um aspects of the community yeah thank you so much for having me this was super fun I wasn't sure what to expect but I I came in and it was a great time and I'll I'll put a shout out for our converge Computing work and flux framework and anyone that's interested in chatting or talking like please find me and let let's talk and you know so much of this work is just about having fun like Whoever has the most fun at the end of the day they win so let's let's do

that together amen girl couldn't have said it better awesome all right we'll see you next time thank you guys have a great day [Music] bye [Music]

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert