RCR: Provisioning: Stateless Vs. Stateful webinar poster

RCR: Provisioning: Stateless Vs. Stateful

Watch Now

Our panel of HPC experts discusses Provisioning: Stateless Vs. Stateful.

Webinar Synopsis:

Speakers:

  • Zane Hamilton, VP Sales, Engineering, CIQ

  • John Hanks (Griznog), HPC Principal Engineer, Chan Zuckerberg Biohub

  • Alan Sill, Managing Director HPC, TTU

  • Gregory Kurtzer, CEO, CIQ

  • Misha Ahmadian, Research Associate, TTU

  • Chris Stackpole, Advanced Clustering Technologies

  • Glen Otero, Director of Scientific Computing Genomics, CIQ


Note: This transcript was created using speech recognition software. While it has been reviewed by human transcribers, it may contain errors.

Full Webinar Transcript:

Zane Hamilton:

Good morning, good afternoon, and good evening, wherever you are. Welcome to another CIQ Research Computing Roundtable. My name is Zane Hamilton. I am the Vice President of Sales Engineering here at CIQ. At CIQ, we are focused on empowering the next generation of software infrastructure, leveraging the capabilities of cloud, hyperscale and HPC. From research to the enterprise, our customers rely on us for the ultimate Rocky Linux, Warewulf, and Apptainer support escalation. We provide deep development capabilities and solutions all delivered in the collaborative spirit of open source. Today's topic, we are actually going to talk about provisioning in HPC, and we are going to talk about stateless versus stateful. And I know this can be a really exciting topic. I know there are a lot of very passionate views on both sides. Let's bring in our panel. Welcome everyone.

Gregory Kurtzer:

Hello everybody.

Zane Hamilton:

How is it going, fellas? I am hoping for an exciting topic. I know there are some passionate views on this, on both sides. I am excited to have this conversation. I am going to let everybody introduce themselves. John, it has been a while. Why don't you introduce yourself?

John Hanks:

I am John Hanks, Griznog, and HPC system for a long time. I don't know how much I should introduce myself since I have been here a lot of times.

Zane Hamilton:

It's true.

John Hanks on Stateless

John Hanks:

My stance on stateless is that in my environments, I target only having three OS installs to disk. Those are two backup Warewulf controllers and DNS DHGP servers, and then one server that acts as a backup should one of those two fail. It is basically just the matching hardware. Then I shoot for everything else in my environment, being provisioned, stateless with the OS and ram.

Zane Hamilton:

Fantastic. Glad I know what side you are on. Alan, if you're still having a little bit of trouble with audio, I am going to go ahead and skip down to Greg. Tell us who you are. We might not know. 

Gregory Kurtzer on Stateless

Gregory Kurtzer:

I have been having this debate for 20 some years now. At this point,  Warewulf has done a lot to really help with stateless installations and stateless cluster management. That was always the first thing that people would always harp on, is it stateless? Is it stateful? How is it working? There was always a lot of convincing that seemed to happen in the early days. This is okay. It is going to be stable, but the operating system's not written a disk. How is it going to be stable because you have memory and memory's good. Those were the conversations, which we would have. I am looking forward to this conversation. I think there are a lot of really great insights as well as pros and cons to doing both. This is going to be a fun one.

Zane Hamilton:

Excellent. Thank you, Greg. Chris, it has been a while. How are you? Welcome back.

Chris Stackpole on Stateless

Chris Stackpole:

Thank you. Staying busy. I am Chris Stackpole. I was working for Advanced Clustering Technologies. We are a system provider, so we will build you out of cluster and we provide a lot of utilities to help manage that cluster. We do both stateful and stateless systems. Personally, I am going to be on the stateful side.

Zane Hamilton:

All right. Keeping track over here, making sure I know who is on what side. We shuffled around. Glen, I think I know what side you are on too, but introduce yourself and give me a hint.

Glen Otero on Stateless

Glen Otero:

Hi Glen Otero, Director of Scientific Computing, Genomics, AI, and machine learning at CIQ. I remember first when Warewulf went stateless. I was really excited about it. I am actually, I am not going to take a side because I think it just, just depends on what you are trying to do. In my last role, my last employer, I found stateless to be great when I was in test and dev, right? I was trying different OS to different OS versions and things like that on compute nodes and applications and testing. But, things in production just did not seem to change all that much. That is where I am at now. It is really great when you have a lot of churn going on and different things you want to try, but in production I do not see as much of an advantage so it depends.

Zane Hamilton:

All right, Misha, welcome back. It is good to see you.

Misha Ahmadian on Stateless

Misha Ahmadian:

Thank you very much. Good to see you all.  Misha Ahmadian from High Performance Computing Center of Texas State University currently we have Warewulf three running on the production cluster, and we just happen to have a new test cluster, which is a small cluster we put together. I was able to install Warewulf 4.3 on it. I am still playing with Warewulf 4 on Warewulf 3 side, you can have both stateful and stateless. We barely use the stateful most of the time. We just shoot the nos in the stateless, state, but I do not know, I am still feeling like if we go fully in the stateless, then I am missing something. Now I know that Greg was able to convince me before that, well, stateless is just enough for you, which I would like to hear more today.

Zane Hamilton:

Excellent. Thank you, Misha. Alan, welcome back. I am pretty sure I know what side you fall onto.

Alan Sill on Stateless

Alan Sill:

Well okay. I will carefully try to split the gap as Misha did. Misha is really leading our efforts here at Texas Tech. We also have some clusters that we run for the other side of my slash here, the NSF funded a cloud and auto computing center where we are running some test clusters aimed at renewable energy settings trying to explore that parameter space of how to run hardware when it is in the middle of cotton fields. Winding back a little of the history, we actually do run the production cluster in a stateful mode. As Misha said, we almost always, if we reboot them, we will choose to reshoot them to fix problems. I am a fence sitter on this. A year ago, I probably would have argued strenuously in favor of stateful, just because I do not like having to dump all the software over the network.

Our old system was, we reinstalled all software after every reshoot, the new scheme, we installed all the user software on centralized resources. We have shifted the burden to those NFS servers, and we have been engaged in a long catch up game with respect to how to provision a cluster when all the software is relying in a central resource. To summarize, in the early days, we used stateful because there was so much software to install in each node. Now that is not the case, but that has really shifted the low elsewhere.

Zane Hamilton:

Excellent. Thank you, Alan. I am going to start off and go back, John, I am going to ask for you to tell us what is stateless? What does that even mean?

What is Stateless

John Hanks:

It is a terrible term. It does not make sense anymore to say, stateful and stateless. What you are really saying is, regardless of where you at your OS, you are going to reinstall it on every boot so that after every reboot, that node will be a fresh install. The idea that you would put the OS in memory is super appealing to me because I like VOS to go really fast. But if you were to have a system that wrote that onto a dispartition, that every node had a dispartition, it would not matter. It is the fact that you are reinstalling from scratch on every boot is what makes it stateless for me. Then the other side of that being a terrible term is I run all my storage servers stateless. Clearly, I am not reinstalling all the data on every boot. Those nodes keep the data alive continuously, the data's always going to be there. That is where the disks are. It is just the OS that we are talking about when we say we do it stateless.

Zane Hamilton:

Thank you. Chris. I am going to let you tell us about stateful. What stateful means.

What is Stateful

Chris Stackpole:

Well, it is interesting because, using that last definition, I am almost on that one. The way I would usually define stateful and stateless is whether or not there is a disk drive to have the OS on. I still highly recommend that your systems are set up such that when you reboot, intentionally reboot, you get a fresh image because that fresh image would be your golden image, which you always know is correct. You can update the image, roll out whatever changes you make sure that is correct. Then every node in the cluster has that golden image. The difference for me on the stateful versus stateless is being able to boot back into the environment when there was an unintentional reboot.  Being able to go back to that OS and have that when things are not going when they are going wrong, being able to go back into that environment as it was.

Zane Hamilton:

Excellent. Thank you, Chris. So coming in from, from an enterprise world and hearing stateful and stateless I had a little bit of a different understanding. I am going to ask this question of Greg. How is stateful and stateless when we talk about HPC provisioning different than when we talk about stateful and stateless in applications?

Stateful and Stateless HPC Provisioning VS Stateful and Stateless In Applications

Gregory Kurtzer:

That is a good question. When we are talking about stateless in high performance computing context, usually not always, but usually what we are talking about specifically is the operating system. And is that operating system written to some form of stateful media? Now, there are, as John articulated, a lot of different ways to pull this all together. I have seen stateless implemented on disk full systems such that even the operating system was written to disk, it was written persistently, but on every single boot it was rewritten. It did not keep that state or that state was not typically persistent, if that makes sense. I have seen it both ways. When we are talking about microservices, when we are talking about applications in application state, it basically is talking about a way of dealing with microservices or services in a way that each one can scale independently.

Thus to do that, you cannot maintain a persistent state for each one of those applications. Each one of those applications have to be able to scale. If you are running an application, let's say in Kubernetes, and all of a sudden Kubernetes needs to scale up 10 more of these, if each one was maintaining state and memory, and that state was required to continue the lineage of that program, well, you cannot actually then just start spinning up more. You have to now deal with other ways of maintaining that state. In high performance computing, we typically limited the conversation of stateful versus stateless to operating system persistence. In many cases, even though this is kind of the well recognized debate, stateless versus stateful, I tend to go back more towards disk full or diskless. With regards to the operating system, and I will give one more point on this, if you can have a stateless operating system that is a disk full system that even manages disk state. You can do that with data. You do not have to do that with the operating system. The operating system could be volatile and every boot it gets rewritten, but each boot has an entry in FS tab to mount up this local file system, put it in the same location. One boot you could be running CentOS, the next boot. You could be running Rocky or Zeus, or Ubuntu, but that data is still exactly the same because it was written persistent to the disk.

Zane Hamilton:

Great. Thank you, Greg. All right, Alan, let's talk about benefits first and then we will come back and talk about the drawbacks of each. What do you see as the benefits of a stateful system in HPC?

Benefits Of A Stateful System

Alan Sill:

All right, so I am going to go back to the late police to scene when we still had a one gigabit control network. At that point, close to a thousand nodes and shooting each one took a lot of time. The six year old cluster, close to six year old cluster we have, has a 10 gigabit control network. The two year old addition has, which is most of our computational power, has a quarter of the nodes of the original setup and 25 gig control network. While the time consideration for shooting those has gone down, we have also, as I said, shifted the pattern to move to the bulk of the load of each of those was actually user software. The time to boot the node, all the cluster is drastically lower than it used to be, but the original motivation was it just took a lot of time to revision each node the way we were doing it before, and we did not want to have to do that every time we rebooted for some especially trivial brief and like changing the image or something like that.

Well, if you change the image, you do have to change it, but I am thinking of some more corrective action that you take on each node you do not necessarily want to reinstall all the software. We just wanted everything that was on disk to stay there. The other thing we did at the newer cluster is we actually cut the disk space in half. We actually do not even have space for all the user software. Logically, you would say these two arguments go against each other if you have shortened the amount of time, amount of things you have to load to each node, why don't you go to stateful? What I am trying to point out was that the items we were trying to preserve on disk had very little to do with the operating system.

If we were clever, we could have done things with partitions and so forth so that we could have had the best of both worlds. Also, I have to say, there is a certain amount of spelunking you have to do in the user support threads. People often had trouble getting one or the other to work. I have actually never been able to discern the pattern. Pretty regularly, you find things in the forums. I can make it work stateless, but I cannot make it work stateful or vice versa. To some degree, it is a random flip of the coin which way you get it to work and then you stick with that. 

Zane Hamilton:

Great. We already have our first question. In my experience, bare-metal compute nodes do not get a reboot very often. Statement, not a question. Very true. All right, I am going to open up and let anybody else add to that of benefits of stateful. Before we talk about benefits. 

Gregory Kurtzer:

I do have a point on that. I have seen certain sites that have done reboots after every job that runs. As a matter of fact, I have seen some sites that have actually even reprovisioned, stateless after each job that runs. Some of these are in classified networks, and as a result that they have to really just manage all of the state and just they want a complete clearing of everything in that operating system before running another job on that system. More typically, I completely agree with the point in terms of how most people run their systems. We actually had stateless systems running at the Department of Energy that I mean in some cases. We could not reboot them just because of very long running jobs and user requirements. Sometimes, I mean, these things were up for I hate to say a year, just because everyone is going to wince at the kernel bugs and vulnerabilities that we might have had, but there have been times in which we have approached that. It does happen under some circumstances. 

Zane Hamilton:

Great. Chris, I think you had something.

Chris Stackpole:

Yes. Coming from my background here where we are dealing with when customers call us, it is my cluster is having a problem and there is a huge difference between my cluster is having a problem and the node just disappears and we have no idea what is going on, versus, oh, I can reboot back into the OS as it was and look at the log files. I can see a lot more information there. We have so many examples of software that really expects the OS to be in a disk form to be able to manage those files versus coming off of a NFS share, which we see a lot for the stateless systems. They have one image that is being mounted somewhere. One of the more recent issues that we saw was with Relic-6, they changed something in network manager and how it works with drag cut.

It was just puking because it expected to be able to write files in Etsy, which were being read only mounted from the file share. We see this a lot with different things. System D has got a lot of those issues as well.  One of the benefits of having it on the OS is who cares if every node is writing changes to Etsy with System D, or being able to capture all the log files of every detail. Especially, when we have seen a job that was causing a broadcast storm across the network it did not really impact the ones that had no US, right, in terms of interrupting the OS ability to function. There is a lot of that localization of having it all on the system so that it can just keep doing what it needs to do.

Now, some of those obviously are going to change because it depends on how you are doing stateless. I have heard people mention where they are moving the entire OS inside of memory, so you are giving up a chunk of your memory for it. However, big that memory is. Some people run really slim of a gig or two just the bare minimum, less other people pack it with every driver that they need and all of their software configuration. Then you are chomping what, 10, 16 gig of memory.  That can add up depending on what kind of jobs you are running, but you solve those problems of having that localized OS too. The exception of saving off log files, but that can be gotten around to doing a remote K dumps and having a remote log server. But now your complexity is going really high up too, because now you have to have servers to capture all that data.

The simplicity of just having an OS that you provision things to on your time schedule for security patches, or whatever and then you just treat that as it is cattle, it is still disposable. You can redestroy it and build it out again on your whim, but it is still isolated for any type of troubleshooting. That really helps somebody like me who is coming in after the fact. I do not know the system, I am just trying to troubleshoot what is going wrong, which tends to be why I like to have those stateful systems.

Zane Hamilton:

John, I think you've had something you wanted to add.

Workarounds for Stateless Problems

John Hanks:

Yes. I would say a lot of those points that Chris just made where stateless is a problem can be worked around fairly straightforwardly. The one about memory though that is a case where I think Greg will back me up on this. I push this about as extreme as anybody could possibly hope to push it. I boot a 16 gigabyte image. OS image and memory to me would be adorable. I boot an image that once it unpacks is at least 32, and I set aside 64 at least gigabytes of memory for that OS to unpack into. All my nodes though have local NVMe drives, which get configured as swap. The minute a job needs any of that OS memory, it swaps out to that NVMe and we do not worry about it anymore.

The memory impact is effectively negligible at runtime because the OS will swap out to the swap space pretty rapidly if there is memory pressure. Because we build our nodes for life sciences, I think my smallest node probably has 256 gigs of ram. This is not a huge problem for us. We would not have very many, usually nodes are 5, 12, or one terabyte, which is the minimum specs we would have for nodes. That problem goes away. Logging and stuff, I realized as he was describing that I also cheat. I almost never do root cause analysis of a failure. I reboot, upgrade, oh, that fixed it, and I move on. I never ask why something failed. I get to avoid that kind of complexity by just ignoring that root cause analysis is an actual thing that people do.

Gregory Kurtzer:

I love that.

Chris Stackpole:

It is harder when you are on a system that has a smaller amount of nodes not a huge budget where all those nodes really matter, and you have one that just will keep rebooting suddenly without any explanation. And you are trying to actually figure out is it hardware problem, is it your job? Those types of things.

Gregory Kurtzer On The Benefits of Stateless

Gregory Kurtzer:

One of the reasons why  I really became a fan of stateless early on was because I was maintaining fairly large clusters with a fairly small crew. It used to be that a single system administrator was expected to be able to maintain 50-100 servers, a hundred if you are really good, right? And these were all kinds of pets back in the pet era. Now, it is, everything is more towards cattle. When we started doing high performance computing, all of a sudden per engineer, this was going up to 500,000 to 750,000 nodes for one engineer. The scale just got massive. Now, how to deal with this, John, your point makes me laugh. It is accurate, right? In many cases, we just do not care about the root cause, you know, conditions of something.

At the same token, one of the things that, to get back to my point, one of the things I love about stateless is all the nodes are the same, or at least all the nodes within a group are the same. They are not only the same, they are identical from the software stack. There is no variation, zero. If you were to have a problem, node 300 is bombing out, it is rebooting, it is seg faulting, kernel panicking, whatever. Well, if the neighbors are working it is hardware period. Take that node out, send it back to the vendor. I have done this and this is how I was able to survive and not lose my mind. Well, anymore than people may already say I have lost and survived in a highly oversubscribed situation where there was just so many nodes I could not possibly maintain it.

I have told this story before, Chris, I do not know if you have heard this one, John. I think you definitely have. There was a time in which we ran an application across the cluster. It was a very performant application, tightly coupled. Every time we ran it, it always failed. It always failed on one node, and that one node would just we would get a seg fault, an application layer seg fault on one node in a stateless cluster. Now this sounds almost unbelievable that it could be a hardware issue, but there was no possibility, right? No other node exhibited this exact symptom. It was just that one node and it was always that one node. If we left that one node out, the job went to completion. But you add that node back in and that was the only application it would do it on, which we were able to at least find and we sent it back to the manufacturer, they sent it back to us and said, it is working just fine.

We sent it back to them again. They sent it back to us. The third time when we sent it over to them, I think it was the third time they said, well, we put everything through its paces. The only thing that we found that was off was the power supply was not quite giving quite enough current at full load. We did replace the power supply, but there is no way that can cause an application layer error. That was it, that was the problem. Power supply was not giving quite enough juice to that one node. Now, if this was not stateless and I was doing the normal process for debugging an operating system, I may never have figured that out just because I am stubborn. I probably would have just sat there and just continued to think there was some sort of weird software error or a memory error or something.

I never would have guessed the power supply. Because I knew all of its neighbors worked perfectly, I knew it had to be, hardware had to be something associated with that hardware. We did try network ports by the way, and other things as well, because we were stumped on this one. That is an example. It is an extreme example though, but it is an example of how having the exact same operating system with no state and no differences between any other nodes is really, really beneficial. And it makes it easy to maintain and to scale.

Misha Ahmadian:

I would like to add on top of that,  it is a good point that you made of stateless using stateless. If we know that all the nodes in the cluster are the same and you start provisioning them all in the same situation. However, we are running into this type of problem here that some of the nodes that just go down. They just hang and you do not have access to their logs after you bring them back up in the stateless fault. For those types of situations, we just do not know what to do. We do not know where the  problem is. Anytime we bring the backup in the stateless, you just see the fresh logs again once you reinstall the operating system. Now, I know that there are some workarounds that you can make the Syslog to write in like NFS shared location or maybe send it through the network.

I believe in a data center with thousands of nodes writing into a central location each node keeps writing to the message, and almost every type of message is every second. It could saturate the network somehow, or also that central location and hard drive that could be an issue. I mean, that is one of the things that stateful is helpful. When we want to debug a node, we put it in a stateful and then reboot it. If the issue happens again, we can look into the logs from that machine.

Gregory Kurtzer:

We have solved this using both k dumps and whatnot, going to a remote system we can catch kernel cores and whatnot. We also do remote Syslog. Your point is really good. If you have thousands of nodes with a lot of Syslog messaging coming through the pipe, I mean, you are adding interrupts, you are adding GIFs, contact switches and whatnot.That does affect performance. Absolutely, there are ways of managing that and to do that more efficiently. But the benefit that you get from that, I think, is well worth the issue there with regards to network bandwidth and even on a stateful system. I would actually suggest doing that. As a matter of fact, at the Department of Energy, we had to do that because we had to collect Syslog for all systems and all nodes and hand that not only to our control server, but actually hand that to the security team. Security got crash logs for every system in the department of DE network that we were on.

Chris Stackpole:

I think that some of these points  you are making are absolutely good and it is what I would do with my systems.The economy of scale here is what I think is being missed because you talk about the customers that I am going to talk to the most often it is a researcher with his graduate student of the semester, not admins, and they have 20 nodes, which are every five, every year they buy five different new nodes and dump the old five. There is no standard between the nodes. It is just however much they can get with the budget that they have been given that semester or that year. When you are talking about that, where you are now having to stand up a whole bunch of different services in order to have things for the remote logging and all of that other aspect, you are asking for a much higher burden when they do not have a dedicated CIS admin. That is not an uncommon use for clusters in HPC is to see sub 50 nodes being used in smaller scales. When you start scaling out to hundreds, oh, absolutely. That is what I have done on most of my systems, even the ones that did not get to a hundred is just to have that. Also, I was the dedicated admin with the skills to build that. I think the stateful makes it a whole lot easier for those users who do not have that skill set or dedicated staff.

Gregory Kurtzer:

Two questions. First, would it help and facilitate if a lot of this came preconfigured from a vendor? Second question, if they do not have the skills necessary to do a remote Syslog, do they have the skills necessary to debug those Syslogs typically?

Chris Stackpole:

That is a great setup to plug my company, Advanced Clustering. Thank you very much for doing that. Yes, I mean, that is what we do. We have an application, a cluster visor. If you are at super computing this year, we are going to have a big demo of it. We pre-build and configure all of it so that when you designate a restart, you can have a fresh new image based off of  a golden image. We pre prep almost all that aspect. Then, we use a cockpit interface to give them a web interface to do most of their administration. When they just need an add a user, it is a couple of button clicks. We do a lot to make sure that it is very simple for our users. I know we are not the only ones that do it, but I mean, I think that is where companies like the one I work for really shine is because we are not targeting these shops that have thousands of nodes because those shops are going to have dedicated admin staff.  We are targeting more of the people that they do have the graduate student of the year.

Alan Sill:

We used to do that and about three years ago we just declared an end to it. If you have money, you can buy resources out of our current system, or you can join a purchasing pool for the next upgrade. We have a standard configuration. We have a few cases where people have noisily lobbied for special circumstances, high memory nodes or long run times, stuff like that. We always find when we accommodate those, the result is wasted CPU cycles. We just do not allow it anymore. As a result, going back to Greg's point the nodes are identical. I mean, we shoot them with the same image whether we keep it on disk or not is our choice. You have almost convinced me to not. We will see how Misha does with his new test cluster.

John Hanks:

I would point out, I manage extremely non homogeneous clusters to the extent that I rarely have more than about six nodes in any given group that will be the same hardware phenotype. I do not think varying hardware as a blocker to doing things statelessly. At least in my case, the benefits I get from it that are not necessarily intuitive, but I take great advantage of would override a lot of additional pain if I had to go through a lot of additional pain. Probably, the worst one of that is I, with some introspection, I realized early on I do not like to write documentation. I loathe writing documentation. I do not take notes. I write nothing down. The fact that everything that is in my environment has to install it boot means that somewhere I have a script that is literal executable documentation for installing that thing and that is checked into a revision control system somewhere. Anybody can come and follow in my footsteps and see there is the script that does that thing. Man, what a good document writer he was. That alone is worth the price of entry to doing everything statelessly.

Gregory Kurtzer:

Can anyone guess what industry Griznog works in based on his language choices? I have never heard phenotype being used for hardware before. That was cool.

John Hanks:

I paid a lot of money for those biology degrees and I am going to use them.

Gregory Kurtzer:

Awesome. I have a story about Glen. When I first met Glen this was about 2000, 2001, Glen. I mean, this was ages ago. I reached out to Glen because he wrote an article in Linux magazine or Linux Journal, and it was printed and text and you can go to the bookstore and get a copy of it. It was talking about Rocksclusters if everybody remembers Rocks. He wrote this whole article on it and I said, I just created this thing called Warewulf and I would really love your take on it. I reached out to Glen about that. I do not know exactly the exact order of operations, but shortly then after Glen comes over to the lab, we sit down, we meet and we talk about all this stuff. I have to say, for 20 some years, I thought I convinced him. I did not know he was still on the fence. I thought he was an advocate. I am just looking at Glen going, oh my gosh, dude, you had me going for 20 years.

Glen Otero:

I have not been on the fence all 20 years. 

Gregory Kurtzer:

Oh, something recent.

Glen Otero:

Something recent. Yeah, yeah.

Gregory Kurtzer:

Oh, okay, let's have it.

Glen Otero on Stateless vs Stateful

Glen Otero:

No, like I said, it was just because a lot of stuff I did, whether it was using Rocks or OSCAR or openMosix or, you know, name your tool, right? It was always testing, validating things. Benchmarking stuff for customers or getting them up and running quickly. It was always just high churn. Stateless was just a godsend at that point. Because otherwise, like everyone was saying, because I was back building clusters using a hundred megabit networks 20 something years ago. Stateless sound was a great idea. Golden image was the currency of the kingdom, but it would just, like everyone has said, it just would take forever to reinstall just a dozen nodes. Like I said recently, particularly in heterogeneous hardware with multiple phenotypes in my last job, stateless and being able to move quickly and not have multiple images that need upgrading and just being lighter weight.

Stateless was great then, but once I tested everything and threw it over to the fence, they thought they would fossilize it and put it into something stateful and just run it in production. That is why it depends. I prefer stateless for OS and memory local disk for scratch. And, like Chris said, depending on your size, what you do at your logs is dependent on that or whether you care like John, for any root causes, just move on. I have not been deceiving you this whole time, Greg. I just want you to know that.

John Hanks:

You actually reminded me of another thing that I take a shortcut on. I realize I get away with a lot because I would never deceive my users into believing I am running anything worthy of the name production. Everything I run is a test system. That is all I run are test systems. I do not have to deal with this. Maybe that is why I can avoid it. I never have to deal with production systems.

Chris Stackpole:

At my previous job, we were researching and we claimed that we did not have five nines, but we had nine fives.

Zane Hamilton:

There you go, Alan.

Alan Sill:

I want to go back to something Chris said. I will preface this by saying, when you watch a computer boot, it is like watching the history of computer science unfold. You watch all these layers, it is, oh, am I enough of a biologist to do this? Probably not. Was it ontogeny recapitulates phylogeny?

Gregory Kurtzer:

One of my favorite sayings ever.

Alan Sill:

Yes, you watch a machine boot and you are just watching lizard forms of computing evolve into chicken forms and into something resembling the current system. A lot of this is being reexamined with open source firmware. A lot of attempts to make this whole process easier and all of it goes away in the cloud, except it does not, because you watch the boot process there and it is still doing this crap. For those of us who have had to flip the toggle switches on a PDP-11, to get it to the point where it would then pick up and read some medium and I can go back further than that if you want me to, but you get the idea of booting is something to be avoided, right?

It is a risky endeavor. You never know if it is going to hit something. That is of course a completely obsolete line of thinking. It is a question of what you mean when you say you are booting. Now, I bring this up because Warewulf specifically has features in the newest versions.  Here I am serving to you Greg, where you can isolate things that you used to think of as part of the operating system install process right after the boot process gets to the point where it can do that. Maybe, we should talk about those for a bit. What are some of the features that you would like to stress, which make Warewulf easier to use? I am thinking of the layers and I have forgotten your terminology actually.

Gregory Kurtzer:

The overlays,

Alan Sill:

 Overlays. Yeah.

Gregory Kurtzer:

There are a few points that were brought up as well that I think that this ties into. We have talked about heterogeneous systems, but in many cases you will have heterogeneous hardware, but you are trying to have similar interfaces. Again, to take this back into biology, it is like different alleles with similar express traits and characteristics in terms of the interfaces and usage of what these HPC systems are doing. There was that, was that good? Hopefully that was good. With Warewulf you could do this, right? You can have different kernel, different kernel features, different boot options and whatnot, as well as different images or different image versions to go out to different nodes and to Alan's point, even just different files or groups of files that are changed per node via templating. You can create templates that would basically say, okay, this node is going to do this or it is going to have this service active, or in this configuration file it will have this section.

 However,  in this other node via a template, you can have a different configuration. You can very specifically express the different features or  pieces that you want to that you want for a particular piece of hardware or a particular role. You could have IO nodes, you could have different forms of storage nodes, you can have parallel storage nodes.  You can have big MEH nodes, you can have GPU nodes, different types of GPU nodes. It is very good at doing all this stuff. At the same time as scaling out a single image to thousands of nodes, you can also subtly change that image or use completely different images for different subsets of nodes. It is very flexible from that perspective. The other thing I just wanted to mention real quick is and I just to come back to this using up, there are multiple ways of doing stateless.

One is something like an NFS route.  Chris, I think you were alluding to that earlier. Then, there is another one that you alluded to as well, which is basically using up RAM using a memory to basically store and run your operating system. That is the way I always prefer to do it. It does have a negative side effect, which is you are consuming expensive RAM where normally you want your applications to use that expensive RAM.  The fixes and the worker, or the work around however you want to look at it is if you have local storage, which I am assuming we do, because otherwise we would not be having that particular debate for that hardware. If you have local storage, you just set up a swap kernel is actually incredibly good at paging temp FS file systems and then paging it onto swap.

Over time, of course you are going to have a little bit of some, not bottlenecks, but some contention through memory as you are currently your operating systems there. Now, we have to go swap it out. Once that happens, over time, what you find is all of the critical pieces of your operating system that get touched often are going to end up staying in RAM in a small footprint. All the other stuff, which is 99% of your operating system, is going to end up heading over to swap, and then basically just residing in storage and on your spinning or flash disk. At the end of the day, you are not actually using that temp space.  As soon as, again, of course as soon as you reboot that system, all that swap space and memory is gone.

You are repurging this is a load on the network. At this point, I think Misha you brought up as well, regarding network, a lot of this just has to do with just the system architecture that you are building. A lot of people that I know that are building these have an ethernet boot and management network and then they have an InfiniBand or some other form of network and they do the same sort of thing with storage, right? They may have local disks or they may not, but they are also going to, they may have NFS for home and they may have Lustre for parallel storage and data storage. The Lustre may be actually communicating over InfiniBand. There are a lot of different ways to slice and dice this, but that ethernet fabric, if you have more than one network, that ethernet fabric really becomes the management plane.

In many cases, you are not actually consuming your data plane to do any sort of either file system transfer or anything. It all depends, again, on the system usage. In many cases, if you are thinking about doing a stateless system, you want to make sure you are building your system, physically building your system, in a way that would actually support that in an optimized way if you have not, stateless systems are pretty much the same architecture as what most people are building. Maybe, some subtle differences, but not much.

Chris Stackpole:

I think that design actually makes a lot of sense for a stateless where you are doing the caching, you have a large fast swap that makes sense. I mean that is very much the way Griznog described his, the NVMes. Then, having that dedicated network for that traffic. Where we have seen that really not go so well is in those like classified where everything that touches a disk has to have encryption at rest and everything else. They just do not put disks in their nodes at all. Those ones are fun for a whole different reason fun.

Gregory Kurtzer:

There is a way, Chris, of doing stateful with no disks. Have you gone that route?

Chris Stackpole:

No. I did look into that a while back. There was just way more network traffic when we were trying to deal with that than what we really had the bandwidth for at the time. It just was not practical for that particular setup.

Gregory Kurtzer:

Gotcha, gotcha.

Chris Stackpole:

But yes, I have looked into that before.

Gregory Kurtzer:

Usually that boggles people's mind. I am surprised nobody stared at me blankly or yelled at me over that one.

Zane Hamilton:

No. It brings up an interesting point, Greg. Because, looking back through enterprise, I have built thousands of web servers. My next question is, where else could this be used? Where else could stateless be used outside of HPC? Because, I think, of all of the servers I have installed and they are all identical and the only thing I am doing is putting a different config for a web server on them. Why in the world am I spending the time to install an operating system on those things every time? Is there somewhere else in the enterprise that would be useful?

Where Can Stateless Be Used Outside of HPC**?**

Chris Stackpole:

I think we are seeing that more, actually. I think we are seeing that because people are spending up the same node type for a Kubernetes cluster or the same node type for Apache web cluster or however they are doing it. It does not necessarily make sense to do cattle type installations when, or I am sorry pets type installation when you treat them more like cattle where you do not really care and whether you are provisioning in an OS via kickstart or something on the fly or whether you are pushing out a golden image or doing a Warewulf stateless really does not matter as much. I think in the enterprise, you are more set up to have things where you have a remote CIS logging where you have a remote K dump, where you have all these tools already built and managed by either somebody else or at least a team of somebody else's.

The other thing is that it makes it a whole lot easier to manage a group of systems like that for things when your security team comes to you and says, we need to have this type of stuff patched. We need to have this level of, whatever your security standard benchmark is, we need to hit that. Well, creating that one type and making sure it is deployed to all the different, and then having something like Ansible Puppet, Chef makes the alterations to put the right config for the right system, all of that. I think we are seeing that happen a lot more on the medium to larger size. Especially, as people who went to cloud and realized, oh man, that is so expensive. We have to bring some of this back in house. But they are already used now to this DevOps type thing where they were just treating these VMs just like cattle. Can we do the same thing locally? I think we are, as we see more of that shift back into localized data centers, some of those DevOps features are going to pull in a stateless OS.

Zane Hamilton:

John, I know you had something,

John Hanks:

I am debating whether I should say it or not. I I think a lot of this I am, anyone who knows me knows I despise large, large IT orgs for many, many reasons. But one of the things I always like to point out is a CIO, when they are looking at their next job, they are not incentivized to do things efficiently. They are incentivized to have a large staff and a large budget so that their next job has a larger staff and a larger budget. A lot of IT orgs evolve into this situation where they are not hiring people who necessarily know what they are doing. They are getting as many bodies on the org chart as they possibly can doing something stateless like what I do where my entire infrastructure is stateless. I try to literally provision everything stateless requires, for better or worse knowing what I am doing.

The idea that I could somehow have this system have 30 or 40 IT staff with their hands in it is completely ludicrous. There is no way this would work in a large org. This is the way this is done, and a lot of stateless setups that I talk to people that do this, it is done because of short staffing, because they do not have a lot of bodies to do all that. IT org enterprise stuff. Even DevOps can do DevOps on a stateless HPC cluster with a couple of people, or you can do DevOps with a giant DevOps team in Kubernetes. Which one is the CIO going to pick? They are going to go with the large team because that makes them to be a more famous CIO. I think that is why you do not see stateless the way we do it in enterprise.

Zane Hamilton:

That is interesting, and I have not thought of it that way. To Chris's point, when you start pulling things back in from the cloud, I still see enterprises having to put tools on them. That generates more of what you are talking about, John. You have another tools team that is putting on all the monitoring the security pieces. You are still, while you may be treating them like cattle, you are still doing so many things to treat them like pets.

John Hanks:

The number one way to make cloud look cost effective is to mismanage your local on-prem resources. Nobody does that better than a large IT org.

Zane Hamilton:

That is fantastic.

Alan Sill:

No cattle or pets, all this sort of decade old terminology. You know, my background is in grids and we have treated worker nodes like insects, right? Provisioning them was out of scope. Now, with cloud, it is in scope and you do have to think about it. I guess I want to mention a couple special cases and maybe Misha can chime in with some more. Typically on an on-prem cluster anyway, and also probably in some cloud situations, you have to have a few special nodes. They may be ones that you would like to have the ability to reshoot, but they do have some special conditions. Login nodes are good examples because they will have definite IP addresses that are unique to each node. Even if they are part of a login pool.

There could be we have not gotten around to provisioning our BeeGFS nodes and our new setup identically that way.  There are nodes that you basically do not want reshot. Even in a situation where your worker nodes are identical, you may have some special purpose nodes that nonetheless you want to manage through Warewulf. Misha, can you think of any other considerations? What is going through your head as you try to make this decision? Because, I am not going to make it, I am going to make you make it.

Warewulf In A Future Stateful Feature

Misha Ahmadian:

I think the new form of state that stateful situation is going to change the design anyway. I do not think that we would use Warewulf the provisioning NFS storage or any type of storage nodes or any node that requires to stay in the same with the same data all the time. I think in the future of Warewulf is just targeting a certain type of nodes, which is going to be the login nodes or worker node. I still think stateless makes sense in our cases, unless we run into some issues that we really need to move to the stateful since there are some workarounds to have the lock files and maybe make the disk partitioning stable across the nodes. These are the things that we would always like to target and make sure that they always have the same partitioning size, have the same size of the swap space and we want, and it is not going to deal too much with the memory, then I think we should be fine on this view run into some other problems.

Now, my question for Greg maybe would be that since we had this conversation here, would you like to consider having this stateful feature in the future release of Warewulf or No?

Gregory Kurtzer:

Believe it or not, this is one of the most common questions that we get. One of the FAQs that we get about Warewulf, which is can, where Warewulf is really good and built around the idea for reasoning stateless, but can we do stateful? We are thinking about how best to do it. If it fits, there is a whole set of Pandora's box that we open as soon as we decide to do that because there is doing it, which actually is not that hard, but then there is doing it right, which is really hard. We are trying to balance and figure out what is the best way of approaching something like that. If we would, most people that we have spoken with, really like the idea of doing the stateless for a couple reasons.

One is something you said, which made me think of this is you mentioned you probably would not do stateless for NFS servers or I will be more general, file servers or some sort of I/O. What if the operating system itself, the container that you are using to provision out was a preconfigured Lustre OSS, right? And let's say we had the entire Lustre stack already pre-created in containers for Warewulf such that you can now go provision, whether it be NFS I/O, whether it be Lustre I/O, GPFS, WEKA, you can choose whatever file system you want, expose that through Warewulf to the rest of the cluster. You can do stateless across the entire thing. To Griznog's point, you can basically take your entire cluster of thousands of nodes and only have one system actually installed on the entire system.

On the entire thing. There is only one hard drive being used for operating system state and management. That is actually exactly how we did it at Berkeley. We had 3,500 ish nodes. I think it was about 3,500. Our Lustre, our I/O nodes, NFS I/O nodes, everything was 100% stateless. We had it was not Warewulf so it was not containers, but we have VNFS images for every piece of our stack, and we can take those VNFS images and just hand them out to any nodes dynamically even, and kind of reshuffle how the cluster works and operates. One point, and I know we are getting near time, so I am going to, Zane, I am going to use this as my closing notes, so do not call on me for closing notes. Stack brought up a point that or a question he was responding to a question about other areas in which this type of cluster could be used, whether it be stateless or just clustering in general.

We have seen Kubernetes now, which is a big one. I remember the first time when Guitar Center and musician friends, anyone who plays instruments, reached out to me and told me that their entire web infrastructure is running on Warewulf. This was obviously almost a decade and a half, two decades ago. This was a while ago, but that was super cool to hear. The one thing that was really surprising to me, really surprising was I always thought that when we think of provisioning, stateless, what we are really doing is imaging, right? We are taking an image of a drive or an image of an operating system. We are pushing that out to compute nodes. We are running that image. Well, what was interesting to me was I thought HPC was the only industry that did this. Turns out we are not. Clouds do this underneath the interfaces that we see. Most clouds that now I have spoken to in terms of underlying architecture as well as very large, very large, IT infrastructures are doing imaging. They are doing it because it is easier and it makes more sense. That may be a follow up conversation as well about imaging versus doing something like config management.

John Hanks:

I would like to just drop a data point in for doing stateless management of ZFS NAS servers. My previous job we had at least a little more than 60 petabytes of ZFS NAS servers scattered between the US, the UK and Europe. All managed stateless from a single image. It worked fantastically well and there is absolutely no way we could have done that with any other method because we only had three admins for the entire setup and that was on top of the cluster and all the other things that we had to do, GP nodes and everything else for all those sites. I would say there is no reason to be afraid to run storage server stateless it works fantastically well.

Zane Hamilton:

Thank you John. And we are up on time. I know we had one comment kind of a question that could lead into a whole other topic, so we will do it very quickly. Bandwidth as a constraint when talking about doing a hybrid, like having an image in a rack and controlling all the network flow in a rack so that you are copying from same rack to server. Thoughts on this? 

Chris Stackpole:

Think it depends on how you want to do that. I think the easiest way is to have one image that is just broadcast out, but you are going to hit a lot of resource constraints because everybody is going to be hitting that one pipe for different reasons. The tool that we have built uses multicast so I can blast out as many nodes as I want all the same image and it does not matter if it is over a one gig pipe or not because everybody is getting the same bits. Then, Rocks used to do it with BitTorrent. Every client every node would install and start up a BitTorrent client for all of its own packages that it would serve. Nobody was ever really impacted by waiting on the head node to give it information. I think that goes to cluster design and cluster design of how you want to move your images around is going to be a whole different topic. Because, it depends on whether you are pushing at a kickstart and doing a fresh install every time you are pushing out a single container image or you are mounting it over NFS. I mean all that is different ways of doing this and bandwidth constraints are going to be different for each of those scenarios.

Gregory Kurtzer:

I want to add to something Chris said, agreed completely by the way. I would add in terms of response to this that bandwidth is a constraint, but typically not, there is not enough data, even when we are provisioning large images. There is just not a big enough requirement to do that segregation by rack. As a matter of fact, I would suggest doing it at a larger scale than just one rack. I know from a management construct, right, it is always nice to have a top of rack or provision out a provisioner on demand or something along those lines. Typically, a rack at the most. I think you are looking at like a hundred nodes if you are super dense. That is nothing for one provisioner. As a matter of fact, you do not really start seeing constraints until you start hitting about 500 nodes.

Those constraints are typically not bandwidth, but as opposed to you get DHCP issues and you get UDP issues on the TFTP streams, that is where you actually see failures first. The first thing to do is to actually manage, at least in my experience, manage your broadcast domains and then as you are managing your broadcast domains, that goes into your architecture of your system. I would actually say 500 is only a problem if you are booting all nodes exactly at the same time. Literally, exactly, at the same time. If you want to stagger them a little bit, you can actually get up to thousands of nodes with one provision or without any problem. Your networking people would probably balk at that.  In terms of, again, broadcast domains. An easy rule of thumb is 1,000 to 1500 nodes per control server and  you can easily go more than that, but that is the number I usually tell people and they usually double it just because we are HPC people.

Zane Hamilton:

Great. Thank you Greg. We are actually over on time. I am not even going to ask for closing remarks guys, sorry. I just want to make sure I give you your time back. We have asked a lot of you guys to be here. We really appreciate it. For those of you watching, please go like and subscribe, but we will see you next week. Thank you guys for coming.

Transcript

foreign [Music] good morning good afternoon and good evening wherever you are welcome to another ciq research Computing Roundtable my name is Zane Hamilton I'm the Vice President of Sales engineering here at ciq at ciq we're focused on empowering the next generation of software infrastructure leveraging the capabilities of cloud hyperscale and HPC from research to the Enterprise our customers rely on us for the ultimate Rocky Linux werewolf and obtainer support escalation we provide deep development capabilities and solutions all delivered in the collaborative Spirit of Open Source so today's topic we're actually going to talk about provisioning in HPC and we're going to talk about stateless versus

stateful and I know this can be a really exciting topic I know there's a lot of very passionate views on both sides so let's bring in our panel welcome everyone hello hello everybody how's it going fellas I'm hoping for an exciting topic I know there are some passionate views on this on both sides so I'm excited to have this conversation but I'm going to let everybody introduce themselves John it's been a while you want to share it yourself all right um John Hanks grisnog and uh HPC system in for a long time I don't know how much I should introduce myself since I've been here

a lot of times but uh it's true my uh stance on stateless is that in my environments I Target only having three OS installs to disk and those are two backup werewolf controllers and DNS DHCP servers and then one server that acts as a backup should one of those two fail it's basically just the matching hardware and then I shoot for everything else in my environment being provisioned statelessly with the OS and RAM fantastic glad I know what side you're on Alan if you're still having a little bit of trouble with audio I'm gonna go ahead and Skip down to Greg tell us who you

are we might not know still on mute there we are who me yeah there you go your turn okay no he said he's having audio problems so we'll we'll go back to Alan um so I've been having this debate for 20 some years now at this point um werewolf has done a lot to really help with stateless installations and stateless uh cluster management um and that was always the first thing that people would always harp on is you know uh is it stateless is it stateful how is it working and there was always a lot of um convincing that seemed to happen in the early

View full transcriptHide full transcript

days um this is okay it's going to be stable but but but the operating system is not written a disk how's it going to be stable because you have memory and memories good so those are the kind of conversations that we would have and uh so I'm looking forward to this conversation I think there's a lot of really great insights as well as pros and cons to doing both uh and so yeah this is going to be a fun one excellent thank you Greg Chris it's been a while how are you welcome back thank you uh staying busy uh I'm Chris stackpool I uh working

for advanced clustering Technologies we are assistant providers so we'll build you out a cluster and we provide a lot of utilities to help manage that cluster um and we do both stateful and stateless systems personally I'm going to be on the stateful side so all right keeping track over here making sure I know who's on what side we kind of shuffled around uh Glenn I think I know what side you're on too but introduce yourself and give me a hint hi uh Glenn Otero director of scientific Computing genomics Ai and machine learning at ciq um I remember first when uh werewolf went stateless I was

really excited about it um and I'm actually I'm on I'm not going to take a side because I think it just just depends on on what you're trying to do uh you know in my last uh in my last role my last employer I I found stateless to be great when I was in in test and Dev right and I was trying you know different OS to different OS versions and stuff like that on compute nodes and applications and testing um but things in production just didn't seem to change all that much and so that's where I'm kind of kind of at now it's it's

really great when you've got a lot of churn going on and and uh different things you want to try but in production um I don't see as much uh of an advantage but you know so it kind of depends all right welcome back it's good to see you thank you very much good to see you all uh Misha madion uh from high performance Computing Center of Texas State University uh currently uh we have werewolf 3 running on the production cluster and we just happen to have a new test cluster which is kind of like a small cluster we put together and I I was able to install werewolf 4.3 on it so I'm still kind of playing with werewolf 4.

um where will three side you can have both stateful and stateless we barely use the state full most of the time we just we shoot the nodes in the stateless state but I don't know I'm still kind of feeling like if we go fully in the state less then I'm missing something now I know that Craig Greg was able to kind of kind of convince me before that well stateless is just enough for you but I would like to hear more today excellent thank you Nisha Alan welcome back I'm pretty sure I know what side you fall onto well um okay so I will carefully

try to split the the Gap as Misha did uh so Misha's really leading our efforts here um at Texas Tech um and we also have some clusters that we run for the other side of my slash here the uh NSF funded a cloud and autonomic Computing Center where we're running some test clusters um aimed at renewable energy settings um trying to explore that parameter space of how to run Hardware when it's in the middle of cotton fields and stuff um so winding back a little of the history we actually do run the the the the production cluster in a state full mode uh but as

Misha says we almost always if we reboot them we'll choose to reshoot them um to fix problems so I am a fence sitter on this um uh a year ago I probably would have argued strenuously in favor of State full just because I don't like having to dump all the software over the network our old system was we reinstalled all software after every shoot the new scheme we install all the user software on centralized resources and so we've shifted the burden to those uh NFS servers and we've been engaged in a long um catch-up game with respect to how to provision a cluster when all

the software is regarding the central research so the to summarize in the early days we we use stateful because there was so much software to install in each node now that's not the case but naturally shifted the load elsewhere excellent thank you Alan so I'm going to start off and kind of go back John I'm going to ask for you to tell us what is stateless what does that even mean it's a terrible term and it doesn't make sense anymore to say stateful and stateless so what you're really saying is regardless of where you're at your OS you're going to reinstall it on every boot

so that after every reboot that node will be a fresh install and you know the the idea that you would put the OS in memory is super appealing to me because I like the OS to go really fast but if you were to have a system that wrote that onto a disk partition you know that every node head of this partition it wouldn't matter it's the fact that you're reinstalling from scratch on every boot is what makes it stateless for me and then the other side of that being a terrible term is I run all my storage servers statelessly so clearly I'm not reinstalling all

the data on every boot so those nodes keep the data live you know they're continuously the data is always going to be there that's where the disks are it's just the OS that we're talking about when we say we we do it statelessly thank you Chris I'm going to let you tell us about State full a stateful mean well it's interesting because uh I I uh using that last definition I'm almost on that one so um the way I'd usually Define stateful and stateless is whether or not there's a disk drive to have the OS on I still highly recommend um that your systems are

set up such that when you do reboot intentionally reboot you get a fresh image um because that Fresh Image would be your golden image that you always know is correct you can update the image roll out whatever changes you make sure that's correct and then every node in the cluster has that golden image um the the difference for me on the stateful versus stateless is being able to boot back into the environment when there was a um unintentional reboot um being able to go back to that OS and have that when things aren't going you know when they're going wrong being able to go back

into that environment as it was excellent thank you Chris so coming in from from an Enterprise world and hearing stateful and stateless I have I had a little bit of a different understanding so I'm going to ask this question Greg how is stateful and stateless when we talk about hpz HPC provisioning different than we talk about stateful and stainless and applications that's a good question um when we're talking about stateless in in high performance Computing context usually not always but usually what we're talking about specifically is the operating system and is that operating system written to some form of stateful media now there's you know

is is John kind of articulated you know there's there's a lot of different ways to kind of pulling this all together and um I've seen stateless implemented on disc full systems such that even the operating system was written to disk It Was Written persistently but on every single boot it was Rewritten so it didn't keep that state or that state wasn't typically persistent if that makes sense um so I've seen it kind of both ways and when we're talking about microservices when we're talking about applications an application state it basically is talking about a way of dealing with microservices or services in a way that

can each one can scale independently and thus to do that you can't maintain a persistent state for each one of those applications so each one of those applications have to be able to scale so if you're running an application let's say in kubernetes and all of a sudden kubernetes is like we need to scale up 10 more of these if each one was maintaining State and memory and and that state was required to continue the lineage of that program well you can't actually then just start spinning up more you have to now deal with other ways of maintaining that state so in high performance Computing

we typically limited the conversation of stateful versus stateless to operating system persistence and in many cases even though this is kind of the well-recognized debate stateless versus stateful um I tend to go back more towards disc full or discless with regards to the operating system and I'll give one more point on this is you can have a stateless operating system that is a disc full system that even manages disk State and you can do that with data you don't have to do that with the operating system so the operating system could be volatile and every boot it gets Rewritten but each boot is has an

entry in FS tab to mount up this this local file system put it in the same same location so one boot you could be running uh Centos the next boot you could be running Rocky um or Seuss or Ubuntu but that data is still exactly the same because it was written persistent to the disk great thank you Greg all right Alan so let's talk about benefits first and then we'll come back and talk about drawbacks of each but what are you what do you see as the benefits of a state full system in HPC all right so I'm going to go back to the late

police decision when we still had a one gigabit uh control Network um and you know at that point close to a thousand nodes and shooting each one uh took a lot of time so the six-year-old cluster close to six year old cluster we have has a 10 gigabit control Network and the two-year-old uh Edition has uh which is most of our computational power has uh a quarter of the nodes of the original setup and uh 25 gig control Network so uh while the time consideration for shooting those has gone down we've also as I said shifted the pattern to move to the bulk of the

of the load of the chin nodes was was actually user software so um the time to to boot the no all the cluster is drastically lower than it used to be but the original motivation was it just took a lot of time to revision each node the way we were doing it before and we didn't want to have to do that every time we rebooted for some especially trivial reason like changing the image or something like that well if you change the image you do have to change it but I'm thinking more you know some corrective action that you take on each node you don't

necessarily want to reinstall all the software so we just wanted everything that was on disk to stay there um the other thing we did in the newer cluster is we actually cut the disk space in half so we we actually don't even have space for the all the user software um and so um you know logically you'd say these two arguments go against each other you know if you've shortened the amount of time amount of stuff you have to load to each node why don't you go to state full but what I'm trying to point out was that the stuff we're trying to preserve on

disk had very little to do with the operating system if we were clever we could have done things with partitions and so forth so so that we could have had the best of both worlds um I also I have to say there's a certain amount of spelunking you have to do in the user support um you know threads uh people often had trouble getting one or the other to work I've actually never been able to discern the pattern but um pretty pretty regularly you find things in the forums you know you know I can make it work stateless but I can't make it work stateful

or vice versa so to some degree it's a random flip of the coin which way you get it to work uh and then you stick with that right okay I've already got our first question in my experience bare metal compute nodes do not get a reboot very often State another question very true all right I'm gonna open up and let anybody else add to that of benefits of State full before we talk about benefits I do have a point on that on that point um uh I I've seen certain sites that have done reboots after every job that runs and as a matter of fact

I've seen some sites that have actually even reprovisioned statelessly after each job that runs some of these are in in classified networks and as a result that they have to really just kind of manage all of the state and and just they want a complete clearing of everything in that operating system before running another job on that on that system but in more typically I completely agree with the point in terms of how most people run their systems um you know we we actually had stateless systems running at Department of energy that um I mean in some cases you know we couldn't reboot them um

uh just because of very long running jobs and user requirements uh and and sometimes I mean these things were up for you know I I hate to say a year um just because everyone's going to wince at the kernel bugs and vulnerabilities that we might have had but there has been times in which we've approached that so it does happen under some certain circumstances okay Chris I think you had something yeah um so coming from the uh you know my background here where we are dealing with um you know when customers call us it's a my cluster's having a problem and um there's a huge

difference between the my cluster is having a problem and the node just disappears and we have no idea what's going on versus oh I can reboot back into the OS as it was and look at the log files um I can see a lot more information there um and we have got so many examples of software that really really expects the OS to be in a um disk form to be able to manage those uh files versus coming off of a NFS share which we see a lot for the um stateless systems is that they've got one image that's being mounted somewhere one of the

more recent issues that we saw was with uh rel86 they changed something in network manager and how it works as a dry cut and it was just puking because it expected to be able to write files in Etsy which were being read only mounted from the file share we see this a lot with different things systemd has got a lot of those issues as well and so you know one of the benefits of having it on the OS is who cares if every node is writing changes to Etsy with systemd or being able to capture all the Vlog files of every detail um you know

especially when um we've seen a job that was causing a broadcast storm across the network um didn't really impact the ones that were had no less right in terms of like interrupting the OS ability to function um so there's a lot of that localization of having it all on the system so that it can just keep doing what it needs to do now some of those obviously are going to change because it depends on how you're doing stateless so I've heard people mention where they're moving the entire OS inside of memory so you're giving up a chunk of your memory for it however big that

memory is um you know some people run really Slim of a gig or two uh just the bare minimum less other people pack it with every driver that they need and all their software configuration and then you're chomping what 10 16 gig of memory and that can add up uh depending on what kind of jobs you're running but you solve those problems of having that localized OS too um the exception of saving off log files but that can be gotten around to doing remote K dumps and having a remote log server but now your complexity is going really high up too because now you have

to have servers to capture all that data so the Simplicity of just having an OS that you provision things to on your time schedule for whatever security patches whatever um and then you just treat that as you know it's cattle it's still disposable you can redestraw it and build it out again on your whim but it's still an isolated for any type of troubleshooting um that really helps somebody like me who's coming in after the fact I don't know the system I'm just trying to troubleshoot what's going wrong um so that tends to be why I like to have those State full systems John I

I think you've had something you wanted to add yeah so I I would say a lot of those points that Chris just made um where stateless is a problem can be worked around fairly straightforwardly um the one about memory though uh that's that's a case where I think Greg will back me up on this I pushed this about extreme as anybody could possibly hope to push it um I boot I I the 16 a 16 gigabyte image OS image in memory to me would be adorable I boot an image that once it unpacks is at least 32 and I set aside uh 64 at least

gigabytes of memory for that OS to unpack into um that all my nodes though have local nvme drives which get configured as swap so the minute a job needs any of that OS memory it swaps out to that nvme and we don't worry about it anymore so the memory the memory impact is effectively negligible at runtime because the OS will swap out to the swap space pretty rapidly if there's memory pressure but because we build our nodes for Life Sciences I think my smallest node probably has 256 gigs of RAM so this is not a huge problem for us and and we wouldn't have very

many usually nodes are 512 or one terabyte you know that's that's kind of the minimum specs we would have for nodes so that problem kind of goes away um logging and stuff uh I I realized as he was describing that I also cheat um I almost never do root cause analysis of a failure I reboot upgrade oh that fixed it and I move on I never ask why something failed so I I get to avoid that kind of complexity by just ignoring that root causing analysis is an actual thing that people do I love that it's harder when you're on a system that has a

smaller amount of nodes um not a huge budget uh where all those nodes really matter and you have one that just will keep rebooting suddenly without any explanation and you're trying to actually figure out is it Hardware problem is it your job um those types of things so one of the reasons why I really became a fan of stateless early on was because I was maintaining fairly large clusters with a fairly small crew and um you know it used to be that a single system administrator was expected to be able to maintain you know 50 to 100 servers 100 if you're really good all right

and these were all kind of pets back in the pet you know pet era and now it's everything's more towards cattle but um when we started doing high performance Computing all of a sudden per engineer this was going up to like 500 750 000 nodes for one engineer and it the scale just got massive now how to deal with this John you you your point makes me laugh it's it's accurate right in many cases we just don't care about root cause you know conditions um of something but at the same token one of the things that to get back to my point one of the

things that I love about stateless is all the nodes are the same or at least all the nodes within a group are the same they're not only the same they're identical from the software stack like there's no variation zero so if you were to have a problem node 300 is is bombing out it's rebooting it's seg faulting kernel panicking whatever well if it's neighbors are working it's Hardware period take that note out send it back to the vendor I've done this and this is how I was able to survive and not lose my mind well any more than people may already say I've I've lost

but um uh and and survive in you know a highly over subscribed situation where there was just so many nodes I couldn't couldn't possibly maintain it and I've told this story before Chris I don't know if you've heard this one John I think you definitely have but you know there was a time in which we ran an application across the cluster it was a very you know performant application tightly coupled and every time we ran it it always failed it always failed on One Note and that one node would just uh we would get a um a seg fault an application layer seg fault on

one node in a stateless cluster now this this sounds almost unbelievable that that could be a hardware issue but there was no there was no possibility right no other node exhibited this exact symptom it was just that one node and it was always that one node and if we left that one node out the job went to completion but you add that node back in and and that was the only application it would do it on that we that we were able to at least find um we sent it back to the manufacturer they sent it back to us and said it's working just fine

we sent it back to them again they send it back to us the third time when we send it over to them I think it was the third time they said well we look we we put everything through its Paces the only thing that we found that was that was off was the power supply was not quite giving quite enough current uh at full load so we did replace the power supply but there's no way that can cause an application layer error it would it was that was a that was the problem power supply was not giving quite enough juice to that one node now

if this wasn't stateless and I was doing normal process for debugging an operating system I may never have figured that out just because I'm stubborn I probably would have just sit there and just continued to think there was some sort of weird software error or memory error or something I never would have guessed power supply but because I knew it's all of its neighbors worked perfectly I knew it had to be Hardware had to be something associated with that Hardware we did try Network ports by the way and and other things as well just because we were you know we were we were stumped on

this one but but that's an example it's kind of an extreme example though but it's an example of how having the exact same operating system with no State um and no differences between any other nodes is really really beneficial and it makes it easy to to maintain and to and to scale so I would like to add on top of that if so it's it's a good point that you made of stateless using stateless if we know that all the nodes in the cluster all this they're the same and you start provisioning them all in the same situation however we are running into this type

of problems here that some of the nodes that just go down they just hang and you don't have an access to their to their logs after you bring them back up in the stateless world so for those type of situations we just don't know what to do we don't know where is the problem and anytime we bring the back up in the stateless you don't you just see the fresh logs again once you reinstall the operating system now I know that there's some work around that you can make the syslogs to write in like a NFS shared location or maybe send it to the through

the network but I believe in in a data center with uh thousands of nodes writing into a central location uh each node keeps writing to the Warlock message and almost every type of messages every second and it kind of it could saturate the network somehow and also that central location and hard drive um that could be an issue I mean that's one thing that stateful is kind of helpful when we want to debug a node we put it in a state full and then reboot it if it if the issue happen again we can look into the logs from that that machine so um we

we have solved this using um uh both you know K dumps and whatnot going to a remote system so we can catch kernel uh cores and whatnot but we also do remote syslog and your point is really good though right if you've got thousands of nodes uh and with with a lot of syslog messaging coming through the pipe I mean you are adding interrupts you're adding jiffies context switches and whatnot um so that that does affect performance um absolutely there are ways of managing that and to do that more efficiently but um the benefit that you get from that I think is is um well

worth the the the issue there with regards to network uh bandwidth and even on a stateful system I would actually suggest to do that as a matter of fact the department of energy we had to do that because we had to collect syslog for all systems and all nodes and hand that not only to our control server but actually hand that to the secure security team so security got crash logs for every system in the department of energy Network that we were on I think that some of these points you're making are absolutely good and it's what I would do with my systems um but

the the economy of scale here is what I think is being missed because um you talk about the the customers that I'm going to talk to the most often it's a researcher with his graduate student of the semester uh not admins and they've got 20 nodes which are every five every year they buy five different new nodes and dump the old five so there is no standard between the the nodes it's just however much they can get with the budget that they've been given that semester or that year um and when you're talking about that where you're now having to stand up a whole bunch

of different Services um in order to have things for the remote logging and all that other aspect you're asking for a much higher burden when they don't have a dedicated sysadmin um and that's not an uncommon use for clusters in HPC is to see sub 50 nodes being used in smaller scales um but yeah when you start scaling out to hundreds oh absolutely like that's what I've done on most of my systems even the ones that didn't get to 100 is just to have that but I'm also was the dedicated admin with the skills to build that um and so I think the the state

full makes it a whole lot easier for those users who don't have that skill set or dedicated staff um two questions uh first would it would it help and facilitate if a lot of this came pre-configured from a vendor first question and the second question is if they don't have the skills necessary to do a remote syslog do they have the skills necessary to debug those syslogs typically so um that's a great setup to plug by company Advanced clustering uh thank you very much for doing that um but yes I mean that's what we do is that um we have an application cluster visor if

you are at Super Computing this year we're going to have a big demo of it um but we pre-build the configure all of it so that when you designate a restart you can have a fresh new image based off of uh you know kind of a golden image we pre-prep almost all that aspect and then we use a cockpit interface to give them a web interface to do most of the administration so that when they just need to add a user it's a couple of button clicks um we do a lot to make sure that it is very simple for our users and I know

we're not the only ones that do it but I mean I think that's where you know companies like the one I work for really shine um is because we aren't targeting these shops that have thousands of nodes because those shops are going to have dedicated admin staff um we're targeting more of the the people that you know they do have the graduate student of the year yeah so we used to do that and about three years ago we just declared an end to it if you have money you can buy resources out to our current system or you can join a purchasing pool for the

next upgrade we have a standard configuration um we have a few cases where people have uh noisily um lobbied for special circumstances High memory nodes or long run times stuff like that um and we um always find when we accommodate those the result is wasted CPU Cycles um so we just don't allow it anymore um and um as a result going back to Greg's Point um the nodes are identical I mean we shoot them with the same image you know whether we keep it on disk or not is our choice you know you've almost convinced me to not but um we'll see how Misha does

with his new test cluster I I would I would point out I I manage extremely uh non-homogenous clusters uh to the extent that I rarely have more than about six nodes in any given group that will be the same Hardware phenotype um so the I don't think varying Hardware is a blocker to doing things statelessly um and at least in my case the benefits I get from it that are not necessarily intuitive but I take great advantage of um would override a lot of additional pain if I had to go through a lot of additional pain probably the worst one of that is uh I

with some introspection I realized early on I do not like to write documentation I loathe writing documentation I don't take notes I I write nothing down so the fact that everything that's in my environment has to install it boot means that somewhere I have a script that is literal executable documentation for installing that thing and that is checked into a revision control system somewhere and anybody can come and follow in my footsteps and see there's the script that does that thing man what a good document writer he was um that alone is worth the price of Entry to doing everything statelessly can anyone guess what

industry that griznog works in based on his language choices I've never heard phenotype being used for Hardware before that was cool I paid a lot of money for those biology degrees and I am going to use them use them awesome I I have a I have a story about Glenn um about Glenn yeah so when I first met Glenn um this was this was literally like what was it 2000 2001 Glenn I mean this was ages ago um and uh I reached out to Glenn because he wrote an article in Linux magazine or Linux journal and it was printed in text and you can go

to the bookstore and get a copy of it and it was talking about um uh rocks clusters if everybody remembers rocks and um he he wrote this whole article on it and I said you know I just created this thing called werewolf and I'd really love your take on it and so I reached out to Glenn about that I don't know exactly the the exact order of operations but um shortly then after Glenn comes over to the lab we sit down we meet and we talk about all this stuff and I gotta say for 20 some years I thought I convinced him I didn't know

he was still on the fence I I thought he was he was an advocate so I'm I'm just like looking at Glenn going oh my gosh dude you had me going for like 20 years I I haven't been on the fence all 20 years oh something recent there's something recent yeah yeah oh okay let's have it no no like I said it was just at uh because a lot of stuff I did whether it was using rocks or Oscar or open musics or you know name your tool right it was always kind of testing validating things benchmarking stuff for for customers or getting them up

and running quickly so it was always just High churn and so stateless was just a godsend at that point because otherwise like everyone was saying you know because I was back building clusters using you know you know 100 megabit you know networks you know 20 something years ago so stateless um it was a great idea um and golden image was kind of was kind of the the uh currency of the Kingdom but it was just like everyone has said it just would take forever to to reinstall just a dozen nodes so um but no just like I said recently particularly in um heterogeneous Hardware uh

with uh you know multiple phenotypes like in my last uh in my last job um you know stateless and being able to you know just move quickly and um not have multiple images that need kind of upgrading just being lighter weight um stateless was great and then but once I tested everything and threw it over to the fence they thought they would fossilize it you know and put it into something stateful uh and just run it in production so that's so that's why I'm kind of like it depends I prefer stateless just meaning OS and memory local disk for for Scratch um and um like

like Chris said depending on your size what you do at your logs is kind of kind of dependent on on that or whether you care like like John you know for any for any for any root causes you know just move on so um so I haven't been deceiving you this whole time Greg I just want you to know that yeah you actually reminded me of another thing that I take a shortcut on I I realize um I get away with a lot because I would never deceive my users into believing I'm running anything worthy of the name production so everything I run is a

test system that's all I run are test systems and so I don't have to deal with this so maybe that's why I can avoid it I I just I never have to deal with production systems in my previous job we were researched and we claimed that we did not have five nines but we had nine fives there you go Allen so you know uh I want to go back to something Chris it's um uh preface this by saying when you watch a computer boot it's like watching the history of computer science unfold you know you know you watch all these layers it's you know like

or am I enough of a biologist to do this probably not was it ontogeny uh recapitulates phylogeny one of my favorite sayings ever yes you you watched a machine Boot and uh you know you're just watching the lizard forms of computing uh you know evolve into chicken forms and and into something resembling the current uh system so uh you know a lot of this is being re-examined um open source firmware a lot of um you know attempts to make this whole process easier and a lot of it goes away in the cloud except it doesn't because you watch the boot process there and it's still

doing this crap so you know for those of us who have had to flip the toggle switches on a PDP 11 to get it to the point where it would then you know pick up and read some medium um and I can go back further than that if you want me to but um you know the idea of booting is uh is something to be avoided right you you know it's a risky Endeavor you never know it's gonna hit some thing so that's uh of course completely obsolete line of thinking and so so it's uh it's it's a question what you mean when you say

you're booting now I bring this up because werewolf specifically um has features in the newest versions and here you know I'm serving to you Greg you know where you can uh isolate things that you used to think of as part of the operating system install process right after the the you know the the boot process gets to the point where it uh it can do that so maybe we should talk about those for a bit so what are some of the features that you would like to stress that that make wearable fee easier to use I'm thinking of the the layers and the I've forgotten

your terminology actually but yeah the overlays uh overlays yeah so there's a few points that were brought up as well um that I think that this ties into we've talked about heterogeneous systems but in many cases you'll have a heterogeneous Hardware but you're trying to you know have similar interfaces so again to take this back into biology it's like different alleles with similar Express traits and characteristics um in terms of the interfaces and usage of what these hvc systems are doing there was that was that good hopefully that was good um but with werewolf you could do this right you can have different kernel different

kernel features different boot options and whatnot as well as different images or different image versions to go out to different nodes and to Allen's Point even just different uh files or groups of files that are um changed per node via VIA templating so you can create templates that would basically say okay this node is going to do this or it's going to have this service active or in this configuration file it'll have this section but in this other node via a template you can have a different configuration so you can very specifically Express the different features or pieces that you want to um that you

want for a particular piece of Hardware or particular role you could have i o nodes you could have different forms of storage nodes you can have parallel storage nodes you can have big mem nodes you can have GPU nodes if different types of GPU nodes so it's very good at kind of doing all this stuff so at the same time as as scaling out a single image to thousands of nodes you can also subtly change that image or use completely different images for different subsets of nodes so it is very flexible from that perspective the other thing I just wanted to mention real quick is

and I just to come back to this using up you know there's multiple ways of doing stateless one is something like an NFS root Chris I think you were talking kind of alluding to that earlier and then there's another one that you alluded to as well which is basically using up uh RAM using up memory to basically store and and run your your operating system that's the way I always prefer to do it but it does have a negative side effect which is you're consuming expensive Ram where normally you want your applications to use that expensive Ram the fixes and the worker or the worker

on however you want to look at it is if you have local storage which I'm assuming we do because otherwise we wouldn't be having that particular debate for that Hardware so if you have local storage you just set up a swap the kernel is actually incredibly good at paging temp FS file systems and then paging it onto swap so over time of course you're going to have a little bit of some you know uh some uh not bottlenecks but some contention through memory as you're currently your operating systems there and now we have to go swap it out but once that happens over time what

you find is all of the critical pieces of your operating system that get touched often are going to end up stay staying in Ram in a small footprint and all the other stuff which is 99 of your operating system is going to end up heading over to swap and then basically just residing in storage in in on your um uh spinning or or flash disk so at the end of the day you're not actually using that temp space um but as soon as again of course as soon as you you reboot that system you know all that swaps face and memory is gone so you're

re-provisioning um this is a load on the network but um uh this is a point I think Misha you brought up as well regarding Network a lot of this just has to do with just the system architecture that you're building uh a lot of people that I know that are building these have a ethernet Boot and management Network and then they have an infiniband or some other form of network and they do the same sort of thing with storage right they may have local disks or they may not but they're also gonna you know they may have NFS for home and they may have luster

for parallel storage and data storage and the luster may be actually communicating over infiniband and so there's a lot of different ways to kind of slice and dice this but that ethernet fabric if you if you have more than one network that ethernet fabric really becomes the management plane so in many cases you're not not actually consuming your data plane to do any sort of you know either file system transfer or anything um so it all depends again just kind of on the system usage and in many cases if you're thinking about doing a stateless system you want to make sure you're building your system

physically building your system in a way that would actually support that in an optimized way if you if you haven't stateless systems are pretty much the same architecture as what most people are building maybe some subtle differences but not much yeah I think that design actually makes a lot of sense for a stateless where you're doing the caching you've got a large fast swap that makes sense I mean that's very much the way grizzdog described his the nvmes um and then having that dedicated Network for that kind of traffic um where we've seen uh that really not go so well is in those like classified

where everything that touches a disk has to have uh encryption at rest and everything else so they just don't put discs in their nodes at all those ones are fun for a whole different reason fun are you doing so there is a way Chris of doing State full with no discs have you gone that route no uh we I did look into that um a while back um and um there was just a way more Network traffic when we were trying to deal with that than what we really had the bandwidth for the at the time um and it just was not practical for that

particular setup gotcha gotcha but yes I have looked into that before it usually that boggles people's mind um I'm surprised nobody nobody stared at me blankly or yelled at me over that one no it brings up an interesting point Greg because looking back through Enterprise I've built thousands of web servers so my next question is where else could this be used or else could stainless be used outside of HPC because I I think of all of the servers I've installed and they're all identical and the only thing I'm doing is putting a different config for a web server on them why in the world am

I spending the time to install an operating system on those things every time so is there somewhere else in the Enterprise would be useful uh John can you take me too we're seeing that more actually I think we're seeing that because people are spending up the same node type for a kubernetes cluster or the same node type for Apache web cluster or you know however they're doing it um and it does not necessarily make uh sense to do cattle type installations when or I'm sorry uh pets type insulation when you treat them more like cattle where you don't really care and whether you're provisioning in

an OS via Kickstart or something on the fly or whether you're pushing out a golden image or you know doing kind of a werewolf stateless really doesn't matter as much and I think in the Enterprise you're more set up to have things where where you've got a remote sys logging where you've got remote kdub where you've got all these tools already built and managed by either somebody else or at least a team of somebody else's um and so um and the other thing on that is it makes it a whole lot easier to manage a group of systems like that for things when your security

team comes to you and says we need to have this type of stuff patched we need to have this level of you know whatever your security standard Benchmark is we need to hit that well creating that one type and making sure it's deployed to all the different and then having something like ansible puppet Chef make the alterations to put the right config for the right system you know all of that and I think we're seeing that happen a lot more on the medium to larger size especially as people who went to cloud and realized oh man that is so expensive we got to bring some

of this back in house but they're already used now to this devops type thing where they were just treating these VMS you know just like cattle can we do the same thing locally um and I think we're as we see more of that shift back into localized data centers some of those devops features are going to pull in such as kind of the stateless uh OS John I know you had something I'm debating whether I should say it or not I think a lot of this uh I I am anyone who knows me knows I despise large I.T orgs uh for many many reasons but

um one of the things I always like to point out is a CIO when they're looking at their next job they're not incentivized to do things efficiently they're incentivized to have a large staff and a large budget so that their next job has a larger staff and a larger budget and so a lot of it orgs evolve into this situation where they're not hiring people who necessarily know what they're doing they're getting as many bodies on the org chart as they possibly can doing something stateless like what I do where my entire infrastructure is stateless I I try to literally provision everything statelessly requires For

Better or Worse knowing what I am doing and the idea that I could somehow have this system have 30 or 40 IT staff with their hands in it is completely ludicrous there's no way this would work in a large or this the way this is done and a lot of stateless setups that I talk to people who do this it's done because of short staffing because they don't have a lot of bodies to do all that it org Enterprise stuff even devops you know you you can do devops on a stateless HPC cluster with a couple of people or you can do devops with a

giant devops team and kubernetes um and which one is the CIA going to pick they're going to go with the large team because that makes them to be a more famous CIO I think that's why you don't be stateless the way we do it in Enterprise and that's interesting and I haven't thought of it that way in kind of to Chris's point when you start pulling things back in from the cloud I still see Enterprises having to put tools on them so that kind of generates more of what you're talking about John you have at another tools team that's putting on all the monitoring the

security pieces so you're still while you may be treating them like cattle they're still doing so many things to treat them like pets the the number one way to make cloud look cost effective is to mismanage your local on-prem resources and nobody does that better than a large it org laughs that's fantastic cattle pets all this uh sort of decade-old terminology you know my background is in grids and we've treated worker nodes like insects right so um provisioning them was out of scope um so now you know with with Cloud it's in scope and you do have to think about it um so I I

guess I want to mention a couple of special cases and maybe Misha can chime in with some more um typically on an on-prem cluster anyway and and also probably in some Cloud situations you have to have a few special modes uh you they may be ones that you would like to you know have the ability to reshoot but they they do have some special conditions uh login nodes are good examples because they'll have definite IP addresses um that are unique to each node even if they're part of a login pool there there could be uh you know we haven't gotten around to provisioning our bgfs

nodes in our new setup identically that way um but you know there's nodes that you basically don't want re-shot um and so even in a situation where your worker nodes are identical you may have some special purpose uh nodes that nonetheless you want to manage through through werewolf um Misha can you think of any other considerations what's going through your head as you try to make this decision because I'm not going to make it I'm going to make you make it yeah I think uh the new form of stateless stateful uh situation is going to change the design anyway so I don't think that we

would use werewolf to provisioning uh NFS storage or any type of storage nodes or any nodes that require to stay uh in the same with the same data all the time so I think in the future of the werewolf is just targeting the the uh a certain type of notes which is going to be the login nodes or or worker nodes um but uh um still I think uh stateless still makes sense uh in our cases unless we run into some issues that we really need to move to the state for um since there are some workarounds to uh have the log files and uh

maybe make the disk partitioning stable across the nodes these are the things that we would always like to Target and make sure that they always have the same partitioning size have the same size of the swap space and we want and it's not going to deal too much with the memory then I think we should be fine now unless we run into some other problems now my question for Greg maybe would be that since we had this conversation here would you like to consider having this stateful feature in the future release of werewolf or no so believe it or not this is one of the

most common questions that we get one of the FAQs that we get about werewolf which is can werewolf is really good and built around the idea for reasoning stateless but can we do stateful um we're thinking about how best to do it um and if if it fits there's a there's a whole set of Pandora's Box that we open as soon as we decide to do that because there's doing it which actually isn't that hard but then there's doing it right which is really hard and so we're trying to kind of balance and figure out what's the best way of of approaching something like that

um if we if we would uh most people that we've spoken with um really like the idea of doing the stateless um for a couple reasons and one is one is something you said which made me think of this is uh you you mentioned like you probably wouldn't do stateless for NFS servers or I'll be more General file servers or or some sort of i o but what if what if the operating system itself the container that you're using to provision out was a pre-configured luster OSS right and let's say we had the entire luster stack already pre-created in containers for werewolf such that you

can now go provision whether it be NFS i o whether it be lusterio gpfs weka you can choose whatever file system you want expose that Drew um through werewolf to the rest of the cluster and you can do stateless across the entire thing into grizznog's point you can basically take your entire cluster of thousands of nodes and only have one system actually installed on the entire system on the entire thing there's only one hard drive being used uh for operating system State and management um that's actually exactly how we did it at Berkeley uh we had 3 500-ish nodes I think it was about 3 500.

uh our luster our i o nodes NFS i o nodes everything was 100 stateless and we had uh it wasn't werewolf 4 so it wasn't containers but we have vnfs images for every piece of our stack and we can take those vnfs images and just hand them out to any nodes dynamically even and kind of re-shuffle how the cluster works and operates and one point and I know we're getting near time so I'm gonna Zane I'm gonna use this as my closing notes so don't call on me for closing notes okay stack brought up a point that um you know or question I was responding

to a question about other areas in which you know this type of cluster could be used you know whether it be stateless or just kind of clustering in general we have seen kubernetes now which is a big one um I remember the first time when uh Guitar Center and Musician's Friend anyone who plays instruments um reached out to me and told me that their entire web infrastructure is running on werewolf this was obviously like almost decade and a half two decades ago this was a while ago but that was super cool to hear but the one thing that was really surprising to me really surprising

was I always thought that when we think of provisioning stateless what we're really doing is Imaging right we're taking an image of a drive or an image of an operating system we're pushing that out to compute nodes we're running that image well what what was interesting to me was I thought HPC was the only industry that did this turns out we're not uh clouds do this underneath the interfaces that we see most clouds that now I've spoken to in terms of underlying architecture as well as very large very large I.T and infrastructures are doing Imaging and they're doing it because it just it's easier and

it makes more sense so that may be a follow-up conversation as well about Imaging versus doing something like config management but I'd like that that is I'd like to just drop a data point in for doing uh stateless management of ZFS Nas servers in my previous job we had at least and maybe a little more than 60 petabytes of ZFS Nas servers scattered between the US the UK and Europe all managed statelessly from a single image and it worked fantastically well and there's absolutely no way we could have done that any other method because we only had three admins for the entire setup and that

was on top of the cluster and all the other things that we had to do GP nodes and everything else for all those sites so I I would say there's no reason to be afraid to run a storage servers statelessly it works fantastically well thank you John and we are up on time I know we had one comment kind of a question that could lead into a whole other topic so we'll do it very quickly so Bandit is as a constraint uh talking about doing a hybrid so like having an image in a rack and controlling all the network flow in a rack so that

you're copying from same rack to server thoughts on I think it depends on you how you want to do that um because uh some you know I think the easiest way is to have one image that's just broadcast out but you're gonna hit a lot of resources uh constraints because everybody's going to be hitting that one pipe for different reasons um but um the tool that we've built uses multicast so I can blast out as many nodes as I want all the same image and it doesn't matter if it's over a one gig pipe or not because it's just everybody's getting the same bits um

and then you know rocks used to do it with a BitTorrent every client every node would install and start up a BitTorrent client for all of its own packages that it would serve and so nobody was ever really impacted by um you know waiting on the head note to give it information um but yeah I think that goes to Cluster design and cluster design of how you want to move your images around is going to be some a whole different topic um because it depends on whether you are pushing out a Kickstart and doing a fresh install every time you're pushing out a single container

image um or you're mounting it over NFS I mean all that is different ways of kind of doing this and um bandwidth constraints are going to be different for each of those scenarios I want to add to something Chris said I agree completely by the way um but I would add in terms of a response to this is bandwidth is it is a constraint but typically not there's not enough data even when we're provisioning large images there's just not a big enough requirement to do that segregation by rack um as a matter of fact I would suggest do it at a larger scale than just

one rack and I know from a management um Constructor right it's always nice to have like a top of Rack or or provision out a provisioner on demand or something along those lines but typically um Iraq you know at the most I mean I think you're looking at like 100 nodes if you're super dense um that's nothing for one provisioner uh as a matter of fact you don't really start seeing constraints until you start hitting about 500 nodes and those constraints are typically not bandwidth but as opposed to you get DHCP issues and you get UDP issues on the tftp streams that's where you actually

see failures first so first thing to do is to actually manage at least in my experience manage your broadcast domains and then as you're managing your broadcast domains that goes into the architecture of your system so I would actually say you know 500 um is only a problem if you're booting all nodes exactly at the same time literally exactly at the same time but if you want to stagger them a little bit you can actually get up to thousands of nodes with one provision or without any problem but your networking people would probably um bucket that in terms of again broadcast domains so a easy

rule of thumb is a thousand to fifteen hundred nodes per control server and you can easily go more than that but that's the number I usually tell people and they usually double it um just because we're HPC people great and thank you Greg we are actually over on time so I'm not even going to ask for closing remarks guys sorry I just want to make sure I give you your time back we've asked a lot of you guys to be here so we really appreciate it for those of you watching please go like And subscribe and we will see you next week thank you guys

for coming [Music]

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert