
Next-Generation Cluster Management: How Warewulf Pro is Transforming HPC Operations
Explore Warewulf Pro in this in-depth on-demand video, featuring the next-generation cluster management platform revolutionizing high performance computing deployments.
In this session, Jonathon Anderson and Rose Stein show how Warewulf Pro transforms traditional cluster deployment from a weeks-long ordeal into a streamlined, days-long process. Built on the proven open source Warewulf project, this comprehensive platform delivers everything you need for full-stack HPC management—from initial node provisioning to ongoing cluster operations at scale.
Whether building your first HPC cluster or migrating from legacy systems, this video will demonstrate how Warewulf Pro delivers the perfect balance of open source innovation and enterprise reliability. Discover why leading global commercial and research institutions trust Warewulf as their cluster management solution.
Transcript
Yeah. So the the I werewolf is is great and werewolf v4 even better in my opinion. Um it is what I tend to say is it is exactly what I in a previous life when I was a cluster assisted admin wish had existed when I was when I was a cluster system. I would have used it all day. I am very excited. Jonathan Anderson, thank you so much for doing all of this wonderful work and I'm so excited for you to be able to show it off. Finally, it is ready to rock. So, >> you know what? It's it's a good day for Werewolf. In fact, >> true.
>> So, so how do we start? What What do we do? >> I I think that first of all, thank you everyone for being here. So, I think we've got 80 people registered to see and talk about and be live. You guys have full access to ask stuff in in the chat. So you be chatting and asking and all of that stuff and they will behind the scenes let us know where the questions are. Um I think probably like super briefly I think everyone kind of knows what werewolf is just thinking it we should give a little intro. So yeah um so we're we're talking about high performance computing today uh which is not necessarily the only use case for werewolf or werewolf pro but uh it's certainly what we're focusing on and in particular um cluster computing.
So the the uh the idea being you have usually some control server and a bunch of compute nodes or some other kind of you know mostly homogeneous node. uh the they can be different kinds of nodes too but you want them all to do a certain kind of work or you want them to all to be configured the same way. Uh so werewolf is a system an open- source communitymaintained and moderated system uh that lets you provision compute nodes or servers or virtual machines if you really want uh statelessly over the network so that when you turn them on they get their configuration and their image their complete operating system from the werewolf server and uh all you have to do is maintain the server and the configuration in one place and every time you turn those nodes on no matter what happened the last time they were on.
They come up into a known good state and uh this werewolf's been around for a long time. I think it was first uh like the first version of like werewolf 0.1 or 1.0 or whatever was a project at LBNL in 2001 by Greg who's you know our our glorious leader at CIQ here. Um and it's it's been a real pleasure and honor to continue that legacy um for Werewolf with in uh its current fourth major iteration. So uh Werewolf Pro, a product that we have recently introduced at CIQ is based on the Werewolf version 4 that has been in development for some number of years, something like four or five years now, maybe more.
View full transcriptHide full transcript
Um >> four or five years. Wait, so where were four is when they did the rewrite into go, right? >> That's right. Yeah. >> And it's been four years, >> I think. So I I I could verify it um real quick, but >> you know what? Because I've been at CIQ for almost three years, but I think for the last three years, I've been saying it's been one year, so you're probably right. >> Oh, yeah. Yeah. So, um I I'm on the GitHub right now and going to the first release. So 4.0.0 was released on January 26th of 2021. Okay. >> So, so thereabouts and of course it was being worked on before that as well.
>> Right. Right. Right. Before that. >> Yeah. So the the I werewolf is is great and werewolf v4 even better in my opinion. Um it is what I tend to say is it is exactly what I in a previous life when I was a cluster assistedman wish had existed when I was when I was a cluster system admin. I would have used it all day. uh but it exists now and I'm glad to be on this side of it. >> Yep. >> Uh and uh but uh it's it's adopted not just kind of on its own in the community but it's a a core part of the the current iteration of open HPC and where previous versions of Werewolf have been a part of Open HBC all the way back to the first release of Open HPC as well which is also exciting.
So um we're glad to have representative from the Open HPC leadership on the werewolf technical steering committee as well. So, uh, continued tight integration there. Uh, but all of these things are tools and and tool boxes. They're very flexible. Uh, and they're ways that a a person, a CIS admin, uh, can build up an HPC cluster and we don't want to do anything to interfere with that. I I personally really like the fact that werewolf remains kind of simple in its intent and purpose. It does one thing. It does provisioning really well. Um, and you can use it to do a number of things, which is why I keep proaricating a bit on whether it's just about HPC.
It's not just about HPC. We have people using werewolf to deploy Kubernetes and proxm mocks apparently and um and we've done storage clusters with it. Really any kind of Linux side cluster computing uh could be a good fit for Werewolf. Uh but right now um what we're trying to do with Werewolf Pro is take some of these use cases starting with the high performance computing use case and particularly the the open HPC style uh high performance computing use case and make that a turnkey experience for people who don't want to have to build everything from scratch on their own. There are all these great tools, but you still have to know that they exist and get them off the shelf and add them together and configure them and put it all together.
Build your images, build your overlays. Um, we want to do that out of the box for you so that you can just get one thing called werewolf pro and plug in the open HPC variant thing in that and have a cluster and uh over time add on additional use cases like Kubernetes like excuse me various storage clusters and not have to deal with that. So that's kind of a a summary of Werewolf and where we're headed with Werewolf Pro and where we are today. I can continue by by showing off kind of the the general demo bit, show off where it is in portal, what it looks like if you get to it.
But anything else we should talk to before we get into that, Rose, or are there questions already? >> Yeah, if there are some questions, you guys just pop them into the chat and we will address them. Um, I I think like are you just prepared to dive right into the pro part or did you kind of want to do a comparison? Um I I it's easiest I think to talk about comparison by showing the pro part. Uh I I don't generally I don't really have any parallel demo or anything ready to go. Um yeah, >> definitely werewolf pro type demo stuff. So yeah, I've got a screen share going.
If we can get that up on the screen, we can start skipping through that. So uh at at its kind of most base level, what is Werewolf Pro? What do I get? If I am a CIQ customer and I get Werewolf Pro, this is technically the thing. Um, we have a a content delivery system that we call portal or the CIQ portal and in there is a product called Werewolf Pro and it is made up of constituent parts. So most obviously uh you get packages for werewolf. So the this is basically the same exact thing that you would get from the community. Uh although uh we are building it locally.
We you get packages from us and we don't necessarily have the same release cadence for this as the community does. The community tries to keep up to about one patch release a month and there's no real guarantee aside from that. Um but if there were a bug that uh one of our customers needed fixed or something that we needed to do outside of the community's release cadence, we could do that here. Uh we haven't had to fork in that way yet. try to keep as close to the community as we can there. Um, we have also added a web interface to Werewolf, which is something we have heard people asking for since I've been here.
Uh, that we we will have a conversation and show off werewolf. That looks great, but does it have a web interface? Well, now it does. And we'll we'll show that off a bit. And then to the point of turnkey uh a turnkey experience, Werewolf is a system for provisioning node images and then configuring them with overlays. So we provide some pre-built node images for common HPC use cases and then overlays that go with those node images that allow you to configure them without having to make any modifications to them directly uh or write those overlays yourself. So things like configuring slurm or configuring PBS, doing some system configuration for time synchronization, um basic stuff now, but as much as Werewolf Pro is a set of features, it's also this delivery pipeline.
Once we have this established, you'll see more node images, more overlays, and you'll just get access to those without needing to build them yourself, uh without having to define them yourself or keep those images up to date. You can just go get the latest version from Werewolf Pro. uh almost uh kind to the end. We also have a separate set of documentation. It's listed as documentation here. This is really more training material. So, we have a training um like a an in-person or uh or remote training that we can give and that that's not included in Werewolf Pro. That's a separate engagement. But as part of what you get with Werewolf Pro, you have access to the curriculum that we have developed and continue to maintain for that training course.
So if you're someone that wants a little bit more guidance on how to set up a cluster, what it looks like, how we intend it to be used, all the way from I don't have any cluster at all and werewolf isn't even installed to I want to build my own images and I want to build my own overlays and how does that work? How does the templating work? Um, you can get that from this curriculum. Uh but you can also engage with us for actual kind of in-person or remote face-to-face training. And then the last thing on our list, oh skipped through the web interface there, is that uh several of our node images include uh Appainer out of the box.
And we're we're very opinionated about the value of containerization in high performance computing. So we wanted to make sure not only that we had Appainainer installed in these images but because it is installed and because we're providing it werewolf pro includes uh support for appainer in the context of the werewolf pro cluster that you have and that support of course extends to the entire stack. So uh anything from uh werewolf and the web interface to the node images that we provide not necessarily all of the software that is in the node images. We pull in a lot of community packages and that's that's not necessarily all in scope for the support but certainly whether the images work as intended and whether the overlays that we provide are are functioning properly.
All of that is backed by uh CIQ support and of course we'll do best effort to help you if you do encounter a community problem inside of one of the images as well. >> Yes, thank you for that additional context because that's a question that I get quite often being on the sales team for this. It's like, well, okay, you have a slurm overlay here. That's cool. Are you providing slurm support? And it's like, well, yes and no, right? >> We'll do our best to help. We'll help you get it configured. Um, not necessarily full stack slurm support. Not necessarily support for every package that is in Open HPC, but you know, our support team's pretty good.
We we do our best to uh help people out when they have problems. So it doesn't mean we can't help but you know there's always a fuzzy line when you're dealing with support in a a broadly community uh maintained ecosystem. >> Right. Exactly. And what I tell people I'm like at least ask. Right. Just don't don't assume the answer is no. Definitely ask. >> You you will almost certainly get more support from us than you expect. Um, but there there's a difference between, you know, our best effort, what we try to do, and our our commitment, which is to the the Werewolf Pro uh components that that we actually build and supply, which is inclusive of Werewolf itself.
Um, sometimes that's a fuzzy line, too. People think, oh, well, that's a community thing, that's an open source thing. Uh, Werewolf Pro explicitly considers werewolf itself within the u the support scope. So, if there's a problem with werewolf, um, we we will fix that. We will do our best to fix that in the community first. and then bring that uh back locally because we don't want to diverge. Um but if we have to, we'll we'll fix things locally first. >> So uh you know, kind of the highlight uh of this description is this web interface. It certainly is a thing uh that is easy to kind of show off and demo.
Let me move some windows around so I can still see people that I'm talking to. >> Yeah. So there actually was a question as well, Jonathan. It looks like Ian, thank you Ian Kaufman has answered the question, but just out of curiosity, I'd love to have you answer it as well. So the question was, is it true that werewolf is used only during the provisioning stage and not for the remaining years when the HPC cluster is in production mode? >> Um, so because expects a cluster to be statelessly provisioned, the nodes are are reprovisioned every time they turn on. So werewolf is uh is used not just at like when the cluster is first installed but anytime you turn a node on and off it's getting its configuration from the werewolf server.
Anytime you make an adjustment to that configuration um you can push that uh configuration change out from the server. Not it's more pulled by the compute nodes from the server. Um the there are various degrees of active communication between the server and the compute node over that compute nodes like booted lifespan like until the next reboot. That's all optional. Um but it certainly is expected that werewolf remains an active part of your cluster management strategy not just on day one when you first install the cluster but ongoing as you manage it. If you um want to get updated packages into your image, those uh updated packages would go in the image on the server and then the nodes would be rebooted into that new image.
>> Thank you for that. And we do have another question as well. This is actually a really interesting one that I Oh, cool. Thank you. We got the question up on the board. I love that. Um thank you, Jordy. Um they say, "Do you have an overlay for setting up Kubernetes, too?" >> So, no. So right now uh fuzball fuzzball too many products rose right now werewolf pro um kind of the 1.0.0 version of werewolf pro is this initial offering of a turnkey HPC experience.
So that includes kind of a base open HPC experience if you just want compute nodes that have open HPC but more typically uh an open PBS based or a slurm based system and and we will demonstrate that the expectation is that in the future uh we will include additional turnkey experiences uh one of which we expect to be a Kubernetes cluster and it it's a little uncertain to me at this point whether that would be the Kubernetes control plane as well being something that's deployed out into the cluster or just workers which would be more of a typical HPC adjacent use case. Um but uh when we provide that that would include overlays that can help you configure the Kubernetes environment as well.
>> Awesome. So that kind of begs the question as well if there's something that like a custom image that someone has paid for separately. Do we then provide that not necessarily inside here where everyone would have access to it but inside like the CIQ portal and then where define what's available to all customers and what's kind of just for one >> that would have to be case by case like let's say we had a customer come in right now and they're like we love Werewolf Pro we want it on Kubernetes today and we will fund that development as some kind of professional services agreement that would probably be something that we would develop but then also make available ailable publicly.
But if it were something specific to a customer, uh I'm actually not certain that we have an answer to that. We would probably provide it through portal, but probably through a customerspecific product that we developed and give them exclusive access to it. >> Yeah, agreed. >> Um >> glad we worked that out live. >> Yes, that's the werewolf part. >> Yeah. >> Um >> that's awesome. >> Uh anything else or or ready to go? >> I think we're good to carry on. Okay, so this is the werewolf web interface. If you're not familiar, it's based on a uh an existing web interface framework that's in Rocky Linux and other enterprise Linux distributions called Cockpit.
So it might look familiar to you. There are other things in it including like general server dashboard things. You can do virtualization in it. It can hook up to KVM. And rather than build a whole new web interface on our own, not just because of the amount of work. In fact, in many ways, that would probably have been easier, uh, but to, you know, decrease the surface that you have to maintain and encourage use of the rest of the, uh, cockpit ecosystem. Uh, we've developed Werewolf Pro's web interface as a page within Cockpit. So uh it's it's just a couple of packages or you install cockpit on your server and also our our cockpit package for for Werewolf Pro and you get this uh this view here and it highlights the the four main entities within a werewolf environment.
So you have the nodes in your environment. Uh you have profiles which are kind of like an abstract node. This allows you to specify and group configuration that you want to apply to multiple nodes. So you can make one profile that is how you want a node to behave and then associate it with multiple nodes. And then on the other side of the fence the images. So these are typically like complete operating system images and the most of the examples that we have in here are rocky Linux based. And then overlays that are kind of like tiny little images uh but they support templates. So you have overlays for specific configuration goals that you have.
Uh so I I mentioned slurm already. There's a, if I can spell correctly, there's a slurm overlay that we have added to werewolf as part of werewolf pro. And using this overlay, you can configure how slurm is uh will behave on your compute image without having to modify that uh image. And then in particular our overlays, like if you've used werewolf before, if you've done this configuration before, you you may have made a configuration overlay. Let's say like a a munchge key. So, Munge is the system that slurm uses to authenticate communication between the nodes in the cluster. You may have um uh just put the munchge key in an overlay because it's static for your environment.
Uh we develop a template that automatically gets that key from the server and uh populates it on your node without you having to provide it here. Perhaps a better example is that uh our slurm damon configuration the agent configuration uh doesn't statically say where the slurm server is. Instead we do that all with uh werewolf template or sorry templates well templates that read from tags and resources. These are are ways to encode configuration in werewolf without having to hardcode it into files. And so we can see that here if I look at this slurm profile that I've defined uh I've said where my slurm controller is where my slurm server is.
So without having to modify the overlay without having to modify the image I can apply that overlay and that image to some nodes and then just give it some configuration for where my server is where my time server is and and go uh and so from here we can demonstrate. I have um I have one node already running this. There's there's two listed here. Uh werewolf node one and werewolf node two. But one of them is real and the other one is uh for demonstration purposes only. Um and werewolf control 2. This is the server I'm on right now. And because I'm in cockpit, I have access to a terminal.
I can do this all here in the browser. Of course, you can SSH into these things and do it there as well, but it's easier to do this than to bounce around between different windows in the screen share for sure. And uh my my compute node as we can see back here is currently configured as an open PBS node. And it takes a little bit of time to import the image and and set all this up in the first place. Uh and it's not very interesting to watch. So we we've kind of done that ahead of time. Uh but we just have this profile uh that or sorry the profile is defined here.
we have this profile that tells it bapbs node and that causes it to use this image and we can demonstrate that it is doing that. Uh I I'll go ahead and remove this output file from a a previous run. Uh but uh if you've used PBS before, we use this q sub command to run my hello.pbs job. And there's a a sleep at the end of this. So it waits a little bit of time just so we can see it in the queue and then see that it is run and and we had this PVS job run on werewolf node one nowish you know. Um so this is you know we have a compute node running in a PVS environment.
But the the kind of crux of this perhaps the the more interesting part of this experiment is that from this interface without having to make substantial changes to anything I can go in here look at this node see what profiles are associated with it say I don't want you to be a PBS node anymore I want you to be a slurm node we'll save that click this button which is something in the future that won't be necessary but if you've been in a werewolf environment before you make the configuration real by uh building the overlays then I clicked it again by accident. Um, and then uh if we can, let's see.
I'm going to switch over to the other window real quick so that um so that you can see the node boot and we won't watch the whole thing, but just to kind of make it a little bit more real. >> That's okay. We do have a couple of questions. Is now a good time or are we right in the middle? Okay, >> I I'll tell you what. Let me pull this up and we'll take questions once the reboot has started. So uh this is the node that as we can see up here is running the PBS image and we're just going to this is a virtual machine of course we're just rebooting it.
So this it kind of goes back to that question from before do you only use werewolf kind of when you provision with the implication being when you first set up the cluster. Um I'm using it now to repurpose this node for something else. But every time this node boots I'm not reprovisioning it intentionally. It's just that's how it is. Every time you turn it on, it fetches its image from the server and says, "What am I today?" Okay, I'm running not just PBS or Slurm, but what version and are there updates? Do I need to get an updated image than what I was running before? Um, it's always fresh, always uh the intended current state.
So, I'll switch back over to the other window and we can take questions while I'm doing that for sure. >> Okay, awesome. Yeah, I definitely um Oh, thanks, uh, Eugene. Sorry, I missed this. Is the web interface for Werewolf only included in the pro? >> That is correct. So, right now, uh to to not put too fine a point on it, the web interface is a sweetener for people that come to us and want the kind of full werewolf experience. Um there are different kind of thresholds for that depending on uh kind of what how the deal gets structured. Um it shouldn't be too hard to get it, but it's not it's not in the community.
the the web interface is part of a CIQ value ad that has not been committed upstream. It is based on um a REST API. That's how we communicate with the werewolf server. And that REST API was uh committed and contributed upstream um and will continue to maintain that. So you can automate your interaction with Werewolf and you know put put it into another interface if you wish. Um but the the cockpit web interface uh as shown not here because it's not sharing now but yeah there this web interface is a werewolf pro specific feature. >> Yeah. And then of course there were some questions about pricing.
I you know we don't want to dodge that but that's going to require a conversation. Uh there's various things to discuss with you and and flush out. So we are very welcome for you can just email me directly rstein sin@ciq.com. You can go to our website as well, ciq.com, and you can fill out anything in there and it will get to us. May >> maybe the one thing we can talk about is how it's it's structured. Sure. In the past, we've tried to uh offer werewolf support before we had kind of feature edition value ads. Um we we offered werewolf support on like a per person model and I personally like where you we would be based on the size of your admin team.
I really like that, but it didn't fit well into people's procurement models for one and then just ended up always being uh oh, is this person on the account or not? And that's not necessarily the person who's having a problem or or uh trying to fix it right now. >> Um so the the base way that we're offering Werewolf Pro today is a more typical like per like the size of your cluster basically, but then what the actual price is, you know, everything's negotiable. We'll figure something out. Um, we want to get this in people's hands. >> Yes. Yes. True. >> Okay. I think we have a couple more questions.
Did you guys want to pop them up there? Okay, great. Thanks, Alex. Does Werewolf have a convenient way to track different hardware profiles or is that something we would need to manage externally and keep track of? >> So, typically we would expect that different hardware profiles would be represented as uh different profiles in here if and when necessary. So an obvious example is I have a bunch of nodes that are all the same and they're CPUbased nodes and I have a bunch of nodes that are all GPUbased nodes. And you might have different images and different overlays to configure those different sets. Um they might be a superset of each other.
So maybe there's like a base profile that all of your nodes get, but then only the GPU nodes get like a special GPU profile. You can do that kind of layering and mixing and matching. uh but there's nothing that's specifically like noticing that the hardware is different and putting different configuration or different images on them automatically. The expectation is that you would have different profiles for wherever you wanted the configuration to diverge and then denote which nodes in your environment get which profiles which can be a many to many uh configuration. >> Awesome. Thank you. And I think we've got one more question from Benode I believe.
Yeah. Hey Benode, thanks for being here. I see that the number and size of your overlays is growing. Is it possible and normal to merge all the overlays to the golden image? >> So it it would not be normal to merge the overlays into the image like the operating system image. Kind of the point is that the overlay is rendered per node and what's in it is different for each node. Uh so whereas the image itself is is the same image for all of the nodes that use that image. So that that's why they're separate. Uh in the past werewolf itself provided what we were referring to or what I've I've started referring to as as monolithic overlays.
One for the system overlay which is only is applied during boot and then one for the runtime overlay. And you can see that distinction here. There's a system overlay and a runtime overlay for each node. Um uh in the past there was one for each of these that we would ship there. I think the system overlay was called uh dubdubdub and netit or werewolf innit and then the runtime overlay was called generic. The problem with this was that we would pack a bunch of functionality into let's say this system overlay the the werewolf init overlay it in order to be applicable to many different operating systems because we have people running it on Debian we have people running it on different versions of uh Rocky and CentOS before it and we have people running it on SUSA.
These all have a really good example is different ways of configuring the network and so we would just do that. we would have all of these network configuration files in the overlay and when you booted up the node it would have all of them and just whichever one is the one that that particular image uses uh it would just use it and ignore the others. Uh even when that was true it was confusing to people. It would be like why do we have all this network configuration all over the system that no one's using. Um, so that that was a point, but then it would be even more confusing within more recent versions of Rocky Linux where there's been a transition from the older IF config style network configuration to the newer network manager style configuration.
And both of those would be on the system and it would be ambiguous which one is actually being used. So people would want to remove things from those monolithic overlays, but that meant modifying them and then those changes would be lost when a werewolf upgrade would come in. So in order to alleviate that, we've broken up the overlays into specific use case purp uh like what they are for, not just these big monolithic overlays, so that you can disable them individually and not have to go modify the overlays. We've separately uh also introduced this concept of a site overlay. So now when you uh when you add an overlay yourself or make a modification, those modifications are treated separately from the overlays that came with Werewolf or came with Werewolf Pro.
Uh you can see these are what we call distribution overlays, but the host overlay here has been modified. I've done so in my test environment. So it's it's identified as a site overlay. We can see a similar behavior here. I can edit this overlay in the web interface. I add another header in here. We'll just say like we can make a a modification and save it. And now we see that our FS tab overlay is a site overlay because we've made modifications to it. uh if we no longer want those modifications, we can throw away that overlay and that will allow the the underlying distribution, the original version to to still remain.
It's still visible. You can't get rid of that. Um so I think that answers that question. I I know there are a couple more, but I I want to highlight one more thing. So we've just demonstrated that this is not just an observational platform or even just a node configuration platform. You can make new overlays and add files to them and modify them. Here you can uh I should have demonstrated this when I had it open, but um you can see what a given overlay template will look like when rendered on a given node, which is kind of nice. This is something you were able to do before, but lots of people didn't know you could do this.
So, one of the purposes of the web interface beyond just being a nice kind of shiny exterior for Werewolf, which is traditionally a command line system, uh is to make things like this more discoverable. Maybe you didn't know that you could see what a template looks like for a given node without deploying it on that node and seeing, but now you can see, oh, it's right here. I I know that I can do this now. And making functionality like this more discoverable is is the actual purpose of the web interface. It's what what we try to have the be the rallying cry of what is good about it beyond just its existence.
So you'll see these little question mark bubbles also kind of on everything that tells you what each thing in the interface does and what it's for and how it interfaces with other parts of the system. You see this as well in the node and profile editor. Every attribute uh I think what werewolf calls a field um of a node is explained what it's for, how it's used, how it interacts with other parts of the platform. So the the goal here is to make it a bit more self-documenting, a bit more discoverable, especially for newcomers to the Werewolf platform, but even for people that have been using it for a while and just didn't know that a new feature was added.
So I'm talking a lot more questions. Yes, we definitely have questions. Um, okay. No, good. Okay. So, are GPU and IB drivers and updates distributed by CIQ or we have to download from Nvidia site? >> We are actively working on the agreements that we hope will allow us to include those in the images. Right now, the images that you get are not uh they're not included in in the so I'll just go over here. So, we have this werewolf node images and the images that we are publishing today. uh is an open HPC base image and an open PBS and a slurm image.
They do not have Melanox or uh Nvidia GPU drivers in them, but we expect in the near future to provide either that they would be included in these images, which is less likely because it makes them bigger, um or that we'll have a variant that includes them probably based on each of these for like an a- Nvidia for each of them. Um, one thing that we might just do regardless, like in the meantime, uh, is include the in the the, um, the IB support that's included with Rocky Linux, we might just go ahead and install that in these base images. Um, you of course can do that yourself once you import these, but again, to further make them a turnkey experience, um, we might just start doing that.
I I need to look and see how big it makes the image. I don't think it's that bad. And then I know that we always get the question about um re you know or a different operating system. People like they look at this are like why is everything rocky? Like I don't have Rocky. I mean a lot of our customers do but not all of them. >> Um so you want to just address that real quick? >> So you you definitely can use re both as the server forwolf pro. So um you know werewolf pro is distributed as a set of of node images that you import into werewolf.
Actually you can do that here. um you can import directly from the CIQ portal into this interface. Um and then packages for installing this interface and werewolf itself on a server. Those packages can be installed and those node images can be used from a rail server if you uh if you desire. Um it gets a little bit more tricky with using a real compute node and it's just because of an issue of entitlement management. So in order to install packages into a rail environment, it needs a a mediated subscription with the rail subscription manager. And uh there are certainly ways that are documented in the community for making >> Werewolf have access to those subscription keys in the node image so that you or when you're building the image um so that it can get packages.
Uh but it just takes a there's a couple more hoops that you have to jump through. Um, one thing that I didn't note though, so we provide images with two bases today. One is uh denoted as rocky Linux and the other which is not uh in this list uh is would be denoted as RLC. So uh CIQ is offering a build or a variant of Rocky Linux that we are willing to diverge from bug forbug uh compatibility uh with upstream re. Um basically we can make bug fixes without them having happened upstream yet. We can actually fix things. Um if you are a customer of RLC or Rocky Linux from CIQ, we uh provide versions of these images that are based on that.
Um, we also, not yet, but in the very near future, will include a build for our RLCH or hardened variant of Rocky Linux, which includes uh pre-stigging and some other security enhancements, including um some kernel level work that we've been doing to to make a a more secure and more compliant version of Rocky out of the box for customers that need that. >> Yes, thank you. That actually begs the question of LTS. What if like one of our customers who actually on this call has the LTS and I mean that's something that we can do for them, right? >> We we can we are not building the LTS images today, but it's reasonably trivial for us to start offering that.
This is it's mostly a matter of just like where did we put it in portal and then give the right customers access to that. as as you start building images like this it becomes a mess of combinotaurics of okay now you build all these images for this LTS variant and that LTS variant what I can say is we have all of this very well automated and uh if as we have requirements for specific bases for these things uh it's it's pretty trivial for us to add those and and give people access to them >> might just take take a few days or a week basically as long as it would take to get a contract signed we would have it ready for Awesome.
Or perhaps it already is signed and we're like, "Oh, yep. Yep. We got it." >> Yep. Yeah, that's not a problem. We would just have to to start automating the build of that specific image, but we're we're already doing it for the others. >> Yeah, that's wonderful. Thank you for that. Um, okay. So, one more question. Let me know. Yep, we're ready for another question. Or am I off the Okay, good. Okay. Can we provision multiple clusters using one werewolf pro? That's an interesting question. >> So it depends on what you mean by cluster. Um in a sense absolutely because from werewolf's perspective um they're they're just compute nodes and you can have different compute nodes run whatever divergent set of configuration you want.
Um there are also people that from a single werewolf server um are provisioning to different networks. that requires a little bit more manual configuration of that's basically you we mentioned here that I had modifications to my host overlay. You would have some modifications you have to make there that that controls the configuration of the werewolf server itself. Um but it is possible to configure across separate divergent networks with one server. In general, I think we would recommend that you do things in a single server or sorry that you that you have a single server for each logical cluster just to make it easier to do upgrades, let's say, and you can upgrade the werewolf server without uh disrupting as many different uh clusters.
But um it's certainly possible. We we would just probably consider that one big cluster. >> Yes. Is maybe the question more like from this control plane from this visual gooey happiness can I have multiple clusters? >> So in theory yes we don't have a good switcher experience today. So this web interface because it communicates to a werewolf server over a rest interface um doesn't need to be uh one for one tied to a given werewolf server. And in fact, you configure how to connect to it here. And if in a brand new install, you would get this and it would ask, hey, what server do you want to connect to?
We're connected to a werewolf server that is running locally, which we would consider the typical experience. Um, if you wanted to uh operate on multiple werewolf servers that are themselves managing multiple werewolf clusters, you could switch to them by just changing the address and credentials if they were different to talk to different servers. If that became a use case that we heard people specifically wanted, uh, we would probably enhance this so that it would save multiple configurations and you could just pick between them. Right now, that's not what we imagine typical. Um, more typical would just be having a copy. Oh, now I'm over here. Uh, that we would have a copy of the werewolf web interface, the werewolf pro web web interface on each werewolf server and connect to it that way.
Cool. Um, thank you for the question. Does the web UI help users create new overlays? >> Uh, I I suppose it depends a little bit on what you mean by help, but you certainly can do it. So, I can make a new overlay here and it is now here and I can add a file to it. And I'll put this at Well, apparently in the past I've done like path to test and I'm I'm always forgetting. This is where it might be good for it to provide a bit more help, but um you have to kind of know what variables are available. Um I think it's node name.
I every time I try to do this live, it's wrong. Yeah, it's it's not node name. >> Yeah. No, it's not. It's like uh it's something else. I remember you doing this. I remember it last time. >> I will go I will go find one. >> Image name. Image name. Node name. I don't know. >> Image name is also there, but we can Oh, it's just ID. Okay. Uh so one oh one thing you can do uh we have for a long time been shipping this debug overlay which is meant mainly as an example that shows all of the variables that that we're adding variables sometimes but uh as of the last time someone looked at this all of the variables that exist are in here and so you can use this as an example to see um what you can do in an overlay.
So, uh, ID is in here, of course. And so, you can get an example from here and then go here to new overlay and put a thing and there see what it would look like rendered here. And if that's not what you want, then you can edit some more and build up your overlay this way. So to you know to a certain extent yes you can certainly build them here and then apply them immediately to nodes in your cluster. >> One thing I have not yet demonstrated that you know I rebooted that node into slur but we never used it. So uh uh we're back over here now and we see that this node is up and ready to go and receive work.
And the example that I have um is we can run this slurm job with srun this time rather than sbatch but just to make it more immediate. And it's not just a slurm job. We're going to run an MPI job and it's a containerized MPI job. So I have this appainer container here. And we can see like this node image doesn't have MPI in it. Uh and all it has is appainer and from that you can run a multi-process multi-node in theory. I have a single node here in my test environment uh MPI job and I didn't have to change anything about this node. I didn't have or about the uh the image and I didn't have to write any overlays or change any overlays.
All I did was tell it with a tag where my slurm server was and it turned on joined the cluster and I think it will run on it. So uh this is why we include aptainer and appainer support as part of werewolf pro is to enable this this level of turnkey experience where you just populate one field turn the node on and you can run parallel work on it. >> Awesome. I think we got a couple more questions. I love the questions. Okay. Does the cockpit plugin need to run on the wolf server or can it run on a remote server? >> Yeah, it can run on a remote server.
Um, one thing that I haven't said explicitly is that the cockpit system uh or the the the Werewolf Pro web interface, it it needs to be running on Enterprise Linux 9. So, it doesn't have to be Rocky, but it does have to be nine because uh since we're integrating into Cockpit, it has to be a a matching version that it expects to integrate with. But, it can be on a separate server. And so if you were running uh a Rocky 8 um cluster server, let's say, and you had Werewolf running on Rocky Linux 8, but you still wanted this interface, you could have a web interface specific server, and then you would just enable access from that server in the API.
There's a little bit of authorization you have to do. By default, it only listens on local host. And then uh you would just put the address of your werewolf server into this configuration box here. So yeah, the the uh the web interface can be on a separate host. >> Yeah. Yeah, I like that too. Um and then the compute nodes though, they don't have to match the head node. Uh >> absolutely. Yeah. And this is a big innovation for Werewolf 4 in particular. In the past and in many similar systems, um there's some tight coupling between the the server, the provisioning server and its configuration and the compute nodes that you're going to provision with it.
For the most part, there is no such coupling for a Werewolf Pro or Werewolf 4 server. Um you can run a Rocky Linux 8 server and provision Rocky Linux 9 compute nodes or vice versa. And the people in the community that are de developing and maintaining the Debian support are able to cross that boundary as well. They can run werewolf on a Debian server and provision Rocky or SUSA images or run a uh Rocky server and provision um Debian or Abuntu images and I in my testing have certainly provisioned SUSA compute images using a Rocky server. So, uh, not not that that would be very common.
Um, but just to demonstrate, there's no coupling there. You can mix and match. Um, the werewolf server doesn't really care what you're provisioning with it. >> Yeah. And the other thing, thanks for reminding us about um needing some kind of enterprise Linux 9 um for, you know, for cockpit also for Werewolf Pro and this you need uh 4.6.3 werewolf. >> Yes. So, and this is mostly because of the the API work that we've been doing. We did ran a tech preview earlier um where we were where we got an early version of this web interface out into our customers hands that required uh werewolf 461 to get the first version of the API and then there were new features that we added to support um mostly the overlay management and editing that we demonstrated uh that appeared in the API in 463.
So, that's what's included in Werewolf Pro. You of course can get earlier versions of the package either from the community or from us. Um, but that wouldn't necessarily be able to provide all the same functionality in the web interface and may break in ways if you aren't using at least 463. >> Cool. And that kind of answers your question, Homer, as well. Um, we kind of already did the tech preview. So, Werewolf Pro is available to CIQ customers. Um, and Werewolf Pro includes this this UI here. Um but the you know you you are still welcome to use werewolf. It is a community version and of course everything underneath is based on the community upstream uh version of werewolf.
So yeah. Okay. One more question. Jess come on in. Hey Matt. Okay. Could you discuss scaling capabilities and concerns? Sometimes we have to power down for data center work. Yep. And then on power up in this model, it can handle uh hundreds of requests simultaneously or slightly staggered. >> Yes. So um the as you might uh not be surprised to hear, we do not have an HPC cluster that we run this on. So we we depend on feedback from our our users on their experiences with this scalability and some partners. What we're hearing right now is that a single server for a few hundred nodes is fine, but if you get up higher than that, um they prefer to run multiple werewolf servers and there are techniques for putting um like a a Damon in front of their werewolf server.
Uh that's something that we want to um include in Werewolf Pro perhaps in the future so that that scalability story is a little bit um a little bit better. May or maybe that'll happen upstream. It depends on how we end up doing it. Um, but right now our recommendation would be in general it should work. It might take longer for them to turn on than you want. Uh, but when I was in that position, I would tend to start nodes up uh, let's say a rack at a time and that would definitely be fine uh, in a werewolf environment today. Um, but it it really depends on the performance of your werewolf server and the size of your image and the performance of your network.
So that there are a lot of variables that go into that. Yeah, I think bottom line we would love to work with you and iterate and get that that communication back and there has been a lot of that as werewolf has been growing especially this last year it's really um kind of exploded in the field. So thanks for >> so qu another question here I I mentioned that there's decoupling between um uh operating system versions and distributions. Um but the question here is is there decoupling for architecture? Um I I actually don't know what like decoupled distress means but I I would take this to oh distros.
Uh so yes I got you. Um yes. So the um the the most common use case here is I have some x86-based compute nodes and I have some ARM based compute nodes. You can do that from a single server. There are a couple of hoops you have to jump through but you can totally do it. So um there is optionally a binary application called dubdub client or the werewolf client that runs out on the compute nodes. If you are using the werewolf client, um you need to provide both an ARM and an x86 version of that binary and then provision the correct one to the correct nodes.
That's something that we intend to improve and make that more transparent and automatic in the future upstream. Uh but uh it today you have to set that up by hand. Um, and then uh you still need to get an ARM image from somewhere. And if you want to manage that image, by default, you won't be able to log into that image and make changes to it from the werewolf server. That's not necessarily what we would recommend doing anyway. We would recommend building the images through some kind of automated pipeline and pulling those into your server. But if you want to manage them interactively on the werewolf server, um we have documented on the community side a way to configure your werewolf server to use uh a processor emulator to allow you to log into a given alternative architecture image and make changes to it.
So all of that is possible. It's not quite a it works immediately out of the box experience, but you can totally do it and we look forward to making that a a more streamlined process in the future. >> Awesome. And what what if I want where it says Werewolf Pro to say Rose Stein is awesome? >> Well, so yeah, less on the customer side and and more on the partner side.
We want to make Werewolf Pro not just something that you know you might get for a cluster that you're administering, but if you're at an HPC integrator or otherwise build clusters for people and you would like custom branding here or want to use Werewolf Pro as a foundation for a larger HPC kind of cohesive experience that you're offering, um, we also have an option and we're offering the ability to to work on a white label version of this and and brand it for your specific use case. So, uh, that's likely a more niche audience that's interested in that, but something I personally am very excited about the opportunity for.
And so, if you are in that position where you build HPC clusters for others and you would like this to be your HPC in uh, infrastructure web interface thing and not just the werewolf pro one, reach out to us and we'll we'll talk. >> Awesome. And what is all that on the on the lefth hand side there where it says system? There's overview, logs, storage, and everything. >> Yeah. So, this this is what I was talking about being like this is all part of co-pilot. So, or or not copilot, cockpit. So, this is the the cockpit overview of this server and the logs that cockpit provides from the server.
So, exactly like you just said, because we're using cockpit as the web interface kind of framework for Werewolf, you are introduced to this broader ecosystem of tools that you can do. I can add accounts to this server through this interface. I can see services. I can see that I haven't updated because I was too chicken to run an update right before the demo. Um, other things and like I mentioned there's a virtualization like a lib interface for this. So you can run virtual machines on it. Lots of things you can do with cockpit. Um, and werewolf pro is just one of them. >> Yeah, that's awesome.
I I there multiple people have been uh pleasantly surprised that we're using cockpit for this and they're like okay good I know that that's familiar to me I like it and let's go with it. So >> it it was it's been a very good and interesting experience. I am a little bit trepidacious about porting this to Rocky Linux 10. uh but this is part of the uh part of the benefit of the enterprise Linux ecosystem that once we do that it'll remain stable for the lifetime of Enterprise Linux 10 too. So uh we'll get that updated soon hopefully before people start asking for it and uh probably won't go back to eight with it but we'll you know we'll have a variant for 10 at some point in the near future.
>> Cool. >> And packages for werewolf for 10 are coming very soon. That's happening in the community. Um, there's already uh builds happening for Rocky Linux 10 and images for Rocky Linux 10 in the community. Um, we don't have those in Werewolf Pro today, but that'll follow quickly after uh those are available in the community. >> Awesome. Um, well, I don't think that there's any more questions. Again, definitely reach out to us. We want to chat with you. We want to talk to you. ciq.com, you fill out anything there, it'll come to us. You can email me directly rstein@ciq.com. Um, just real quickly, is there any fun things coming?
I mean, it's already the end of August, which I can't even believe right now, Jonathan. So, SC is coming up. Do you want to do a little shout out to what we're doing at SC? >> I mean, we're we're showing off developments in Fuzball. We've got cool stuff to show off from the the workflow catalog that we've been building in Fuzball, as well as performance improvements in the scheduler and provisioner system. Those have been highlights of our development right now to to make fuzzball more kind of scalable to much larger environments and job counts uh for like high throughput computing stuff. We'll be showing off werewolf pro as well.
Um we'll be talking a lot about Rocky Linux hardened there. There's been a lot of interest there as well as another variant that we're building Rocky Linux AI Rocky Linux from CIQ for AI. um uh that makes like diverges from the Enterprise Linux kernel and provides a newer kernel with better performance and better uh compatibility with more accelerators. Um lots of stuff. >> Good show. >> Yeah. And on the community side, uh the community werewolf community is having their first wug the werewolf user group. I'm super excited and grateful to you, Rose, for kind of carrying the mantle on that. This is something I've been wanting to see in the community for quite a while.
So, it's exciting that it has started. >> It's very exciting. We've already got 30 people registered for it. So, it's going to be on Zoom September 10th, Wednesday, September 10th at 10 a.m. Pacific time. Um, so hopefully that works for you and again, give me a shout if you want an invite to that. I think we actually have one last question. >> One last question. >> One last one. >> Nvidia recently made BCM free. What do you tell community to say with Open HPC plus Werewolf? >> I that's news to me. I will need to do some Googling about that here in a moment. But uh I I would say like Open HPC and Werewolf is an open community- based platform that we think is is good not just for the people using it but for the people administering systems with it that they can participate in those communities.
Um perhaps they're going all the way with BCM as well. That would really be news. Uh I will also say um like I've I've never been a Bright uh administrator. I've never worked in that environment. Uh but my the story I get from people who have is that when you administer a Bright system, you are not a an HPC admin or a Linux admin. You are a Bright admin and you do things the Bright way. And uh it it works in just that way. And and an example that people give me is they can't upgrade slurm because the version of slurm that they run is the version of slurm that comes uh with bright cluster manager.
Uh that is not the case in werewolf. Of course we will provide a specific version of slurm and a node image that we consider a turnkey experience. But these are all very flexible tools that we provide support for even when you're bringing your own node image, bringing your own overlays and configuring it yourself. These are all things that have turnkey experiences but also provide configurability anywhere along the stack. And that is a a core tenant of the werewolf design ethos ethos and the werewolf pro design ethos that we want things to be easy for people that want the easy thing but flexible and powerful for people that want to do whatever they want to do with it.
So, we think Werewolf is kind of the right way to do that and we'll continue to improve it. But yeah, this is this is news to me. I'm going to have to go do some reading. >> Yeah. Yeah, that's awesome. We love free. Who doesn't love free? >> Who doesn't love free things? >> Cool. Well, this was a wonderful webinar. We are like just pushing time, man. That was great. We did a full 60 minutes. Jonathan, thank you so much for your time. Thank you everybody who popped in and spent time with us. Thank you for your questions and we look forward to supporting you in the future.
>> Have a wonderful day.
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.