Turnkey Solutions With Rocky Linux, Warewulf, and Apptainer
Webinar Synopsis:
Speakers:
-
Zane Hamilton- Vice President of Sales Engineering, CIQ
-
Gregory Kurtzer- Founder of Rocky Linux, Singularity/Apptainer, Warewulf, CentOS, and CEO of CIQ
-
Robert Adolph- Chief Product Officer & Co-Founder, CIQ
Note: This transcript was created using speech recognition software. While it has been reviewed by human transcribers, it may contain errors.
Full Webinar Transcript:
Zane Hamilton:
Good morning. Good afternoon. Good evening. Wherever you are, welcome. For those of you who have been with us before, welcome back. And for those who have not, we appreciate you joining us to learn more about CIQ. Like our video and subscribe to stay current with us. This week we will be talking about traditional HPC. I have Greg Kurtzer and Robert Adolph, our founders, with us today to have that conversation. Let's go ahead and get started and talk about what traditional HPC means.
Traditional HPC [00:28]
Gregory Kurtzer:
I can jump in and take that. Hi everybody. So I'm old. So when I talk about traditional HPC, usually that means something very specific to me, with regards to tightly coupled parallel applications and an architecture that is designed to optimally benefit applications of that type. Over the years, HPC has expanded to say anything that's CPU-bound, memory-bound, or IO-bound. Anything hitting the hardware directly is starting to be considered high performance. But when I talk about traditional HPC, I'm going back to the foundation and back to the roots. When we talk about how to build an architecture infrastructure to support a traditional HPC application, we usually start talking about an architecture we've used for 28 years called Beowulf. The Beowulf architecture has been an absolute foundation and godsend to how we build traditional HPC resources. It is a flat monolithic architecture, but it's an architecture that's given us the ability, for decades now, to fairly easily scale these big HPC architectures and resources. So traditional HPC means something very specific to me.
Zane Hamilton:
Excellent. Robert, anything to add to traditional HPC?
Robert Adolph:
I'd add that this is a very big focus of some of our solutions. It is an area where our team has a lot of expertise. When CIQ first started, our goal was to bring that type of expertise together with cloud and hyperscale and to enable the next generation of what that could be based on those experiences in both, bringing together the knowledge and know-how of how to run and do parallelization at scale that HPC has been doing for 20 years with the innovation that’s in hyperscale and cloud. That is a driving force behind what we've been trying to do for our whole stack, not just our traditional HPC stack.
Gregory Kurtzer:
A really important facet of Beowulf's architecture is its specification of how to build these systems. One aspect that is super important to talk about is that Beowulf is about commodity computing. This is about getting systems that everybody has access to, turning them in, and creating a cluster—through these systems, with this commodity hardware, creating a cluster using open source software, predominantly to build these HPC resources and systems. Again, I just want to reiterate how important and critical Beowulf has been for decades of research, science, and supporting these sorts of applications.
Zane Hamilton:
You mentioned something interesting. I hear the term open source get thrown around a lot. I've been around open source for a long time, but I know you two are extremely passionate about open source, maybe more so than many others. One of the reasons I'm excited to be here is because you guys focus very much on open source in the community. I want to understand, from your standpoint, what should companies be doing? How should they be going about releasing open source software?
What should companies be doing?
Gregory Kurtzer:
Oh my goodness. That's such an awesome question. I'm a biochemist by degree, and what pivoted my career was because I got so enamored and excited about two things: Linux and open source. Going way back in time now, in the mid-nineties, we were able to create genetic tools for genomics using open source software. We were able to download Linux off of the internet for free, which, as you can imagine, back then was just mind-boggling to be able to download your operating system, put it on a bunch of floppy discs, and start installing Linux on a piece of hardware you just got from Fry’s, and actually build a scientific tool out of that.
Open source had an extraordinarily profound effect on how I see solving problems. I want to be very clear on this, it is a very beneficial and collaborative development model, but it's not a sales tool. It's not a marketing tool. It's not even a commercial asset, or at least it shouldn't be considered as such. The problem is that it is being considered as such by a number of organizations out there. And as a result, I've seen this many times and I’ve even been a part of this, where companies have decided to either adopt and control an open source project or keep it within their corporate umbrella. And what we see is that an open-source project becomes an asset to that company, and that company can choose to leverage that asset for business purposes. The result of this, to be blunt, is that they're holding it hostage. They're holding that open source project and community hostage to benefit that company. One of the primary benefits of open source is the collaborative development model and the freedom and stability it gives its users. When this is controlled and owned by a single company or entity, it loses a lot of that stability. We've seen a number of applications and open source communities that were pivoted in the last five-plus years to benefit a corporate agenda, and it hurts the open source community. It hurts the open-source community quite a bit. I spoke with multiple companies who said they feel better paying for their software and not being open source because they're worried that an open source project is just going to go away if a company decides to change it, and they don't want that. They would rather know that there's some stability, which is very odd because open source should be about that stability.
The way to get an open source project to be super stable is not to have a single company, even a really big company, standing behind that open source project and community while controlling and running it. The way to truly get stability from an open source project is to have a wide variety of organizations, individuals, teams, and people all standing behind this and not having it controlled by a single organization. It was a hard lesson that I've personally had to learn. I've learned through mistakes and, moving forward, I will never support a company, that I am either part of or running, to own and control and hold hostage an open source project. So everything that we are doing from an open source perspective, we are not only putting it back out into the community, but we don't even want to control it; we don't want to own it; we don't want to hold it hostage. We want that community to thrive. We want other organizations, individuals, and people to have control over that group and that project. We believe in that wholeheartedly, and so every open source project at this point that we have released and we have been part of, we don't control. We've done that, and we're going to talk a little about that more, as this webinar continues. I think that's very, very important, and to be blunt, I challenge other companies to do the same. If you love open source, set it free.
Zane Hamilton:
It's an important topic to talk about because many open source projects that were fantastic have gone away. You have to pivot away from that product or that set of tools, and it can become very challenging for the enterprise or even HPC to do that.
Robert Adolph:
As a product manager concerned about solving problems for customers, I find it to be one of the best ways for us to discover new use cases or see new ways to solve different problems. People are so unique and creative that they come up and try to do things themselves. For me, it's a great way to stay in touch with exactly what our customers at CIQ want, while not holding it back and controlling it. Now we can shape it and do what we think we can go and do with our resources and customers. To me, it's the ultimate way to understand the community at large and serve them even better.
Gregory Kurtzer:
Robert, you've talked about how every startup and company is always thinking about product market fit and how to get that product market fit. If you have an open source project, it is to give that open source project to the community, and the community will shape and define what that project means and solves for them. You get instant product-market fit when you have adoption of an open source project that a particular company doesn't own. It's so valuable in so many different ways.
Zane Hamilton:
I appreciate that perspective, guys. I agree. Back on CIQ, how does CIQ enable that traditional HPC stack? I know that there are several different pieces of that, but what does CIQ provide in that regard?
CIQ and the Traditional HPC Stack [12:08]
Gregory Kurtzer:
Over my career, I've had a fantastic opportunity to create various capabilities that have been widely utilized in high performance computing. CIQ is basically pulling the innovations and capabilities that I've created among other people, much smarter than I, that have been able to add value to HPC centers, to users, to people that are building these systems, to commercial HPC. We're now pulling all of those together to create a stack, a traditional high performance computing stack. Three main pieces we have been very close to in terms of innovation have been provisioning, a base-operating system running the resource, and containers. If you take those three major pieces, you have the bulk of a high performance computing system. Obviously, there are more components to this, like batch scheduler file systems, but this allows us to start creating a turnkey solution for organizations from bare-metal going all the way up through the stack, and that's where we're leveraging. I'd like to take a moment and go through and talk about some of these three major areas that we're contributing to that high performance computing stack.
Provisioning is a really big piece of this. To describe what provisioning is, if somebody wants to build an HPC cluster or even any kind of cluster, it's a whole group of computer systems, and they will be built in such a way that they will work together to collaboratively solve a software problem.
Now, if you have a fairly small cluster, let's say in the tens of nodes, it might be possible to go through and handle them like any other server, whether you are using config management or installing them by hand and then managing them by hand. But very quickly, when we talk about these clusters, we are actually scaling way beyond the tens and into hundreds if not thousands of compute resources. When we do that, all of a sudden, the traditional way that you think about maintaining a server needs to be pivoted a little bit because it's just too many systems to do in the traditional sense. So we've been doing provisioning of operating systems, resources, and management of that cluster through a variety of tools.
I created one called Warewulf back in 2001. And Warewulf brings some very unique features to the ecosystem. One of the biggest features–probably the biggest one that it changed and pivoted–was the ability to have stateless compute images or nodes. What this means is, let's say you have a rack of a hundred computers. You roll this rack up into the data center, you give it power, you hook it up into the network, and you literally can start turning on nodes. You don't have to install them. You don't have to configure any of them, and you can turn them all on, just power up the whole rack. We've been able to build high-performance computing systems from the software perspective. From the moment the hardware's all set up, hundreds of nodes, and I'd even go as far as to say thousands of nodes, we can have the whole operating system built and ready to run across thousands of computers in an hour or two.
And the whole thing is ready now for production. And that sort of ease of use and scalability, put together, were some of the defining characteristics and problems we were trying to solve initially with Warewulf back in 2001. Warewulf has done that very well, and as a result, it is still used today. Now we've got a number of different organizations that are also part of the project, including OpenHPC, which, if you're familiar with high performance computing, is a Linux Foundation project to help organizations create turnkey solutions for traditional HPC. Warewulf is the provisioning subsystem for this.
One of the ways that Warewulf solves this is it uses a node image. We used to call this the virtual node file system. Imagine a chroot sitting in a directory on a control node somewhere, and that becomes the representation for every node that will boot. And that's how we've done Warewulf for literally over 20 years. With Warewulf version four, the main pivot has been instead of using this chroot, this virtual node file system, let’s instead pivot to using containers. Let's leverage the container ecosystem to be able to blast a particular container out to thousands of nodes and run that container directly on bare-metal. And that's one of the major areas in which Warewulf has just been so incredibly helpful and useful to everyone doing computing.
The last numbers I heard, we had like 50% of the top 500 systems running Warewulf. That's a huge number in terms of how many people are running Warewulf. There are a lot of systems out there running Warewulf. Imagine if you can use the same tooling you're using currently, whether they're user application containers or enterprise-focused containers; you have a CI/CD pipeline; you have all your DevSecOps; processes in place to validate those images. Imagine if you can now import those into Warewulf and then blast it out to thousands of compute nodes or tens of compute nodes. It makes it that easy and utilizes things like preconfigured containers, just like we use containers today. We can have a container for OpenHPC; we can have containers for Slurm, LSF, Grid Engine, and PDS, just for the scheduling side. We can include InfiniBand, GPUs, Lustre support, FPGAs, and everything into these images. Then we can have these custom images that are targeted to do different things. People can now import those images, just like you do with a Docker image, import that image into Warewulf, and then say, we want to leverage, we want to build a Slurm cluster that's running Lustre or GPFS, and we need it to do X, Y, and Z.
Zane Hamilton:
Yeah, that's fantastic.
Gregory Kurtzer:
Boom. Done.
Zane Hamilton:
Yeah. I want to also point out, you've said it a couple of times, but starting in 2001, that's 21 years of experience with this thing running in different environments. I think that's why it's grown so much, and it's exciting that it's continuing to grow and gain functionality and popularity. Every time I get on the phone, it seems like I'm talking to people using this today, so it's exciting.
Gregory Kurtzer:
You’re just jabbing at my age right now. I got it. I got it.
Zane Hamilton:
The main takeaway from that is you're talking about the fact that it’s stateless, centralized, flexible, and simple to integrate. You can do a lot of preconfiguration and add to that CI/CD pipeline. I think that’s fantastic. From what I've been able to tell, it's also really easy to get it up and running, right?
Gregory Kurtzer:
Yep. Robert, you came off mute. Do you want to jump in?
Where to take Warewulf [21:05]
Robert Adolph:
That's a lot of discussion around where we're going to take Warewulf and what's next. Obviously, the community drives a lot of that. There are a lot of great folks in the community, representing a lot of large organizations. What I would say is obviously a very simple GUI type of graphical interface to help people that are newer at these solutions to get them up and running extremely efficiently and effectively. This is an extremely lightweight solution, so it's designed specifically for this, potentially adding some state fault capability and some APIs to make it more extensible. Then lastly, I would say that potentially utilizing this platform as a large-scale imaging-as-a-service solution for customers of Rocky Linux, customers of HPC. So beyond just HPC, we think we could utilize this lightweight capability at scale.
Gregory Kurtzer:
That's a really good point, Robert. As we've been talking to enterprises, it seems as though you can generalize how people manage their resources based on the size of the organization. On one side, we have people doing physical installs, doing system administration almost by hand on a number of systems. Then we've got this middle ground, where the majority of organizations seem to sit, which is they're using some sort of Kickstart and configuration management-based solution to manage that infrastructure at a larger scale and to be able to do management of the config management system, using even something like Git. So all the configs, everything is in Git, which is then being propagated via configuration management. And you can manage a large number of resources.
But what I didn't realize until fairly recently was the other extreme. Again, one stream is everything by hand, going all the way past configuration management. And looking at it from this extreme, what we’re seeing is organizations with very large clouds, some hyperscalers, they don't do configuration management. Well, they do to some extent, but for the most part, they're doing imaging. So every system is being blasted out as an image. What we've been doing in high performance computing is extraordinarily similar to what's been happening at the extreme scale of cloud and enterprise. We're handling and managing this from an imaging perspective and provisioning out those images of operating systems out to those resources, whatever those resources may be or require, in an extraordinarily scalable and efficient way. That's one of the areas that I just learned about fairly recently that that’s how they're doing it.
And I thought it was cool how similar we've been doing it from an HPC perspective as well as the extreme scale. One of the points that Robert just made is: what if we can change that barrier of entry using something like Warewulf to help organizations do even more imaging-as-a-service across their organization? And maybe even couple that with configuration management, to where we’re actually supporting even smaller enterprises to mid-size and large enterprises, all using these same sort of methods and models.
Zane Hamilton:
Yeah. I think it's really interesting. We've been talking about this from an enterprise perspective for many years of getting to being able to, as soon as it's time for something new, tear it down, build new, and build it up. You got the blue-green deployments, and it's all fine and good in theory, but rarely does it work the way you want it to. So it's exciting to see that there are people making progress down that path. I mean, I feel like this has been ten-plus years in the enterprise, we talked about trying to get to that level, and you don't see many people making that kind of progress towards it. I'm excited to see where this goes as well. From an HPC perspective, we have Warewulf to do provisioning, but you have to provision something. So that next layer up is Rocky Linux. Let's talk about Rocky for a minute.
Rocky Linux [25:43]
Gregory Kurtzer:
I'm going to start it from the perspective of the relationship between CIQ and Rocky. And just to hit the elephant in the room, I'm the founder of Rocky and also the founder and CEO of CIQ. One of the decisions that I made, again, looking back at our open source stance, our goal is not to own, control, or hold any of our open source capabilities hostage. We didn't want to put this in CIQ. We wanted this to be a separate project and a separate organization. That's why I created the Rocky Enterprise Software Foundation, but CIQ has been along with Rocky and with RESF as a founding supporter and partner of the organization.
From day one, CIQ, myself, and others at the company have been backing what we've been doing with Rocky. Our goal, again, is not to do anything from a corporate agenda perspective. We want to facilitate wherever this community wants to go. Our goal is if anybody needs help, we want to earn that business. We want to prove that business, and that's our model, and that's it. That's the relationship between the two. That leads to what we could offer in terms of the operating system portion of this stack. We add various benefits on top of the base Rocky Linux.
And the other thing I'll mention is CentOS is about, in our approximation, all the metrics that we've seen, about 20-25% of all enterprise resources, cloud instances, and containers worldwide. That's a big number. Now that has just been end-of-lifed. Rocky is stepping up to become that replacement, that alternative to what CentOS has been and how CentOS has been trusted for so many years. We're hoping to even do it better in the sense that there's not going to be a single company that controls it. Rocky's going to be completely neutral from that perspective, and we're going to have lots of companies being part of that. But CIQ is one of those companies helping, sponsoring, and partnering to bring value to customers, so we're there to help. We're here to ensure that if anybody has any questions or problems, we can offer that; we can become that enterprise-level production support arm and offer some services.
Now, we already talked a little bit about imaging as a service, but there's also something else that we've had a lot of really great feedback from the community and customers, which is repositories–"repos-as-a-service.” Imagine if a particular organization or site can be able to easily overlay their custom changes to the operating system and make that operating system "theirs" for their organization. All their customizations, everything that they need for that organization, even configurations being pushed out and managed via a repo. That repo has a bunch of code around it that allows organizations to easily add packages into this, easily build packages through that, do cryptographic signing, validation, and everything that they need as part of that repo. You can think of it almost like what a satellite provides, but there's no licensing per node. It gives you the ability to manage, update, and control what an organization's Linux needs are without ever having to worry about licensing costs or anything along those lines. So, with repo-as- a-service, we're getting a huge amount of positive feedback.
We're also focusing a lot on FIPS and other security accreditations, which really has been huge for federal. What surprised me was that many companies out there prefer to know that those accreditations have been met and that they can trust the operating system, so we are moving forward with FIPS. We have a timeline. If anyone's interested in that timeline, please let us know. And we also can block security exploits that use buffer overflows as the vector of attack. Imagine if all past, present, and future exploits, using buffer overflows, were mitigated again. We have the ability to bring that level of security directly into not only the base operating system but also containers, containers running Rocky as that base image. What we're doing from the high performance computing perspective is to add all of this into that traditional HPC stack. You can imagine that coupled with Warewulf and then coupled with containers, that creates an amazing base platform for doing traditional HPC.
Robert Adolph:
I'll do the opposite this time. Usually, I take Greg back from the open source to CIQ and the solutions we're doing. I'm going to go the other way, since he told everybody about our repo-as-a-service solution. The other thing I would highly recommend is there is a special interest group in the Rocky community that is specific for high performance computing. Every engineer at our company has a deep understanding of high performance computing. There are tons of people in the Rocky Enterprise Software Foundation that have a deep high performance computing background. Rocky's being tested on all sorts of high performance computing apps, use cases, storage models, etc. I can tell you that folks in the testing community are high performance computing experts. I would highly suggest getting into that community aspect and looking around and seeing all the resources there. You'll be massively surprised by Rocky Linux's compatibility with high performance computing applications and solutions.
Gregory Kurtzer:
I completely forgot the point I was getting to when talking about CentOS being such a big piece of the market at 20-25% of all enterprise resources. In HPC, it's closer to 70-75%, so Rocky not only has a lot of interest in enterprise hyperscale and cloud, but the interest in high performance computing is massive. If you look at the people that are part of that Rocky community, two people leading the testing group are both from HPC centers. We definitely have a large amount of interest and background with the high performance computing section of the industry. Again, to reiterate what Robert just said, we have a SIG HPC, a special interest group specifically designed to further enable high performance computing on Rocky. If this is something that people are interested in, we will post the link to our Mattermost, which is our open source Slack that you can join. Then, you can search for the SIG HPC channel, and please join. We'd love to have you as part of the Rocky community.
Zane Hamilton:
Excellent. Thanks, Greg. So now we have provisioning, we have an operating system; now it's time for people to start deploying apps. I mean, nowadays, everybody wants to talk about containers and CI/CD. I know you alluded to it a little bit earlier with Apptainer, but tell me about Apptainer. Where does it fit in this model? We've got the base; we've got the OS; now what?
Apptainer and Batch Scheduling [34:37]
Gregory Kurtzer:
Apptainer started life as Singularity, and a lot of people within the high performance computing world know Singularity very well. It has been the absolute dominant market share leader of containers within HPC. If you're not in HPC, it wouldn't surprise me if you've never heard of Singularity. But if you are part of the HPC community, I'd almost put money on that everybody has heard of this and potentially even been using it. It has become such a prevalent and critical piece of high performance computing infrastructure. This is because when containers were starting to evolve and gain traction, publicity, and interest in the enterprise and cloud space, the methods for dealing with the run time for containers were very root-centric. It didn't allow very well for a Beowulf-style architecture that may have hundreds or thousands of users logged into them, with none having any privilege or access above what they would have as just a generic, non-privileged user. To give them access to a container subsystem that was all running on root would've defied all of our security aspects. That would be a security exploit if a user has the ability to control the host operating system. That's exactly what Docker would've given them at that stage. Plus, there are several other facets, such as how we make good use of GPUs, InfiniBand, how do we do things like MPI, which is the message passing interface, the library, then high performance computing that we use to enable parallelization of compute jobs across node boundaries? We can have a hundred nodes all running the same application and communicating between those nodes in a highly optimal manner on whatever network backplane you have. We had to figure out how to do all of that.
How do we integrate with a batch scheduling system that controls all the resources of that HPC system? Long story short, we needed to develop a way of integrating containers into an HPC system that fit with the HPC architecture and use cases. And that was what Singularity did. Within just a matter of months, Singularity went from just an idea to it was probably installed on the majority of big HPC systems that support containers, in a matter of months, not years. The uptake was incredibly fast. It solved a critical problem; it was highly optimized for high performance computing use cases and architectures, it supports things like MPI, Infiniband, GPUs, file systems, and any batch scheduler you want. It also allowed any non-privileged regular or average ordinary users to leverage any container they wanted.
It took off, and we've done things differently from the traditional container ecosystem. Like we have single file-based containers that we can cryptographically sign and encrypt. We bring even a level of security that differentiates the container supply chain further than anything else to my knowledge that's even out there today. Singularity made a bunch of great strides. We wanted to ensure that Singularity stays open source and in the community, for the long haul. We decided, in collaboration with the Linux Foundation, to bring Singularity into the Linux Foundation and allow the Linux Foundation to take over ownership and control of the project. Their one request, and we were on the same page, was that they asked us to rename the project, so they could properly protect the copyrights and trademark. We renamed the project, went out to the community, and said we need to rename the project if we're going to move this into the Linux Foundation. And we gave a few suggestions, other suggestions were offered and Apptainer was the leading vote. At this point, Singularity had been renamed to Apptainer, we have a thriving community around this, and we're part of the Linux Foundation. One of the goals of moving to the Linux Foundation was how do we better cross-pollinate between the high performance computing world and the enterprise and cloud world in terms of container and container infrastructure. We've been focusing on that cross-pollination and leveraging those resources on both sides.
Zane Hamilton:
Excellent. Robert, I know you're very passionate about the security aspect of Apptainer or Singularity and how that single file can have a cryptographic signature. Do you want to dive into that a little more?
Security with Apptainer and Singularity [40:04]
Robert Adolph:
Yeah. I still remember the first time Greg explained to me in detail what he'd created. I said, “Wait, let me get this straight, Greg. Basically, you're saying that I can take my application and move it to any cloud, any Linux operating system, and run on any piece of hardware or anywhere I want in the world and not be completely locked into any cloud or anywhere, or any specific piece of hardware ever again?” He said, “Yeah.” “So you're also saying that it's immutable and perfect for every kind of research possible out there?” “Yes.” “And it's immutable? And reproducible? You literally just solved like 25 problems that I've heard from customers over and over again.” I heard everybody talking about serverless, and I asked, “Why would I want serverless when I can move my apps anywhere I want in the world? Why would I want to be locked into anything?” To me, this was huge.
Greg said something really important: it’s software supply chain security aspects. We can cryptographically sign every step of that process. We can also sign for security signatures, deployment signatures, and DevOps signatures. Then with our new platform, we can ensure that it doesn't get run anywhere unless it has all the correct signatures. This was a basis for many of the future steps we took. It's all because of those 25 problems Greg solved with that one, as Greg says, Pandora's box that he opened. That led us to create the next generation of high performance computing that we've also been working on. We'll talk about that next week.
Zane Hamilton:
Excellent. How does CIQ bring all of this together? What does it mean for a customer?
What does this mean for the customer? [42:13]
Gregory Kurtzer:
Some feedback that we've gotten is, and this is both from vendors and customers, is that they want a turnkey solution. They want to know that they can come to a single location, get the entire foundation of what they want to build, and know that that can be supported and integrated easily. We looked at all these different pieces we've created, and were continuously maintaining and contributing to, and said, let's put these all together into that turnkey solution and build a traditional HPC stack, specifically to be easy to deploy, highly scalable, very flexible, and become that base foundation for HPC centers, HPC consumers, and even enterprises that are looking at doing compute-focused workflows. Let's simplify that, democratize that in such a way that anybody can go and do this.
What I love about how we approached this was that we were not going to make any architecture statements or compatibility options. I don't know if somebody wants to use an NVIDIA GPU or an AMD GPU; I don't know if they want to use an FPGA, I don't know if they want to use Slurm, if they want to use LSF, Grid Engine, or PBS; I've got no idea what a particular customer is going to want to use. I want to provide them with that infrastructure substrate so they can build the tools they need to create and solve their problems. And that’s what we’ve created. We've created a turnkey solution that can go and build up an OpenHPC cluster or a custom cluster, or however they want to build it, easily, efficiently, and scalably.
It's very simple, flexible, capable, scalable, and secure. That was our goal, to bring all of that together. We have some organizations leveraging us for an à la carte stack. In a manner of speaking, some customers are strictly for Singularity and Apptainer; other customers focus on Warewulf and or Rocky–we have many customers on Rocky right now–it could be à la carte as well. However an organization wants to deal with this, we are here to support the organization.
A little bit off topic, but I think it's super important to mention we are offering both capabilities and licensed values on top of open source pieces as well as offering services and support. Most organizations are doing what the legacy model has been for services and support. When we started thinking about how we would do this, we decided to completely rethink and be willing to throw out every idea that's not working and not beneficial to the customer. We decided to make sure that everything we are doing is about the enablement of the customer. If it's not enabling the customer, if it is a value to us and not the customer, then that's not a win for us. I can give you an absolute easy example of this. Many organizations will charge for support based on core, socket, or even nodes. Some organizations have to use nodes to get it through their procurement department. Most people don't like it. They don't want to count the number of systems or resources they are paying for support.
Who does that model help? It helps the company that's offering the support, not the customer. If a customer came to us and demanded per-node pricing, we'd do it, but we wanted to ensure that we were offering something better. Even site licenses can be difficult, so we decided to rethink this completely. Every person, organization, and team that we have shared this model with, even the teams that said, ”We never want support because we have a great team, don't need the help and don't need the insurance policy,” even those companies are joining us because they love the model and they see how we can now help and add value to their team as people, as individuals. If anyone is curious about exactly what that looks like, reach out to us, and we'd be happy to walk you through it.
Zane Hamilton:
Absolutely. If you have any questions, go ahead and start posting those. Robert, you were about to say something.
Robert Adolph:
I take it one step further. Greg told you about some of our design principles. One of them was being modular, in how we approached this, and having the building blocks, but also making it all work better together over time. If you were a Dell customer of Dell hardware, you could utilize their Omnia stack, and utilize Warewulf, Rocky, and Apptainer. If you're an HP customer, you could potentially utilize Rocky, then Apptainer, and maybe keep some of their secret sauce they utilize. If you're a customer of Penguin, you can utilize Rocky; you can utilize Apptainer, in a traditional stack like many folks do today. That's just an example of some folks you could partner with. With Supermicro, you can do the exact same thing: you could use our whole stack or you can use pieces of it that make sense to facilitate the additions to what you're doing.
So definitely work with your OEMs. One of the overriding principles that Greg and I talked about a lot about all our open source projects, all of our projects in general, were to have every cloud, every OEM, every component of a server, and every ISV that could run on any of those environments, as partners, sponsors, or folks we work with on a regular basis, and it's not by name. If there is an issue with a system, I want to be able to call that OEM from our support group and explain to them what's going on. We do not want to make it so the customer isn't getting served. We want people to get their problems solved. If a driver needs to be executed on or something needs to be done, we're there to help them do it. Then the last thing I'd say on containers: obviously there are a lot of libraries and OS capabilities that get built inside of there; we have a full staff and team to enable the entire stack that we're talking about and give your applications speed and life and flexibility, while giving your end users and your admin complete control. That's what we were doing when designing all of our software and potentially in our open source solutions as well.
Zane Hamilton:
Perfect. There aren't any questions. I just heard Greg's reminder for his next meeting pop up. I know you're busy, so I want to respect your time and appreciate you joining. Again, we ran out of time and didn't get a Greg story. Sorry, maybe next time we will push for that. Or maybe we can force Robert to tell a story. We'll see where we get. Anyway, we appreciate you guys joining again. Go like and subscribe, and we'll see you probably next week.
Gregory Kurtzer:
Thank you. I can guarantee there'll be a story next week.
Zane Hamilton:
Awesome. Thanks, guys.
Gregory Kurtzer:
Bye, everybody.
Transcript
good morning good afternoon good evening wherever you are welcome for those of you who have been with us before welcome back and for those who have not we appreciate you joining us to learn more about ciq go ahead and like and subscribe to this to stay current with us this week we're going to be talking about traditional hpc and i have greg kurtzer and robert adolph our founders with us today to have that conversation so let's go and get started and just talk about what exactly does traditional hpc mean i can jump in and take that um hi everybody uh so i'm i'm i'm old
so when i talk about traditional hpc usually that means something kind of very specific to me uh with regards to tightly coupled parallel application and a architecture that is designed to optimally benefit applications of that type uh now hpc over the over the many years has expanded to basically say anything that's cpu bound or memory bound or i o bound anything that's really hitting the hardware directly is starting to be considered high performance and but when i when i talk about traditional hpc i'm going to kind of go back to uh the foundation back to the roots and um when we talk about how to
build an infrastructure and architecture to support a traditional hpc application we usually start talking about an architecture that's um that we've been using now for literally like 28 years called the beowulf and the beowulf architecture has been a absolute just foundation and godsent to how we build traditional hpc resources it is a it's a flat monolithic architecture but it's an architecture that's given us the ability for decades now to be able to scale um and fairly easily scale these big um these big these big hpc um architectures and resources um so yeah it's a traditional hpc means something kind of very specific to me excellent
robert everything to add to traditional hpc no the only thing i'd add is that um you know this is a very big focus of some of our solutions and it is an area that uh obviously our team has a lot of expertise and our goal as ciq when we started out was to bring that type of expertise together with cloud and hyperscale and to enable the next generation of what that could be based on those experiences in both so bringing together the the knowledge and the know-how of how to run and do parallelization at scale that uh hbc has been doing for 20 years with
View full transcriptHide full transcript
the innovation that's in hyperscale and cloud is a driving force of what we've been uh trying to do for our whole stack so not just for our traditional hpc stack so real quick zane if you don't mind me jumping back in oh absolutely a really important facet of the beowulf architecture is is basically it's a specification of how to build these systems and one aspect about it that's i think super important to talk about is the beowulf is about commodity computing this is about getting systems that everybody has access to and cr turning them in and creating a cluster [Music] with these systems with this
commodity hardware creating a cluster using open source software predominantly to build these uh these hpc resources and systems and again i just want to reiterate how important and critical that beowulf has been for decades of research science and um and and supporting these sorts of applications you mentioned something that's interesting i hear the term open source get thrown around a lot i've been around open source for a long time as well but i know you two are extremely passionate about open source maybe more so than a lot of others out there uh fro from what i've been around and one of the reasons i'm excited
to be here is because you guys focus very much on open source and the community and i want to understand from your standpoint what what should companies be doing how should they be going about releasing open source software oh my goodness that's such an awesome question uh and yeah i mean i've i've been so i'm a biochemist by degree and i have got really enamored and what caused me to pivot my career was because i got so enamored and so excited about two things linux and open source going way back in time now in the mid 90s we were able to create genetic tools or
tools for doing genomics using open source you know software we were able to down linux download linux off of the internet for free which as you can imagine back then was just it it was it was mind-boggling to be able to download your operating system put it on a bunch of floppy disks and start installing linux on a piece of hardware that you just got from fry's or or from wherever and and actually build a scientific tool out of that so open source has had an extraordinarily profound effect in terms of how i see solving problems uh it is um and i want to be
very clear on this it's a very beneficial and collaborative uh development model but it's not a sales tool it's not a marketing tool it's not even a commercial asset or at least it shouldn't be considered as such problem is is that it is being considered as such by a number of organizations out there and as a result i've seen this many times and i've even been part of this where uh companies have decided to either adopt and control uh an open source project or not um or or keep it within their umbrella their corporate umbrella and what we see is that open source project becomes
a a asset to that company and that company can choose to leverage that asset for business purposes the result of that is well to be blunt they're holding it hostage and they're holding that open source project and community hostage to benefit that community that company now the way to get and this is really critical because one of the primary benefits of open source is not only the collaborative development model but it's the freedom and stability that it gives to its users and when this is being controlled and and owned in a matter of speaking by a single company by an entity it it it loses
a lot of that stability and we've seen a number of applications and open source communities that were pivoted in the last you know i'd say five plus years uh to benefit a corporate agenda and it hurts the open source community it hurts the open source community quite a bit i spoke with multiple companies who actually said you know a couple of them even gone so far as to say they actually feel better paying for their software and it not being open source because they're kind of worried that that open source project is just going to go away if a company decides to change it and
and they don't want that they'd rather be rather know that there's some stability which is very odd because open source should be about that stability so the way to get a project an open source project to be super stable is not to have a single company even a really big company standing behind that open source project and community and controlling it and running it the way to truly get stability from an open source project is to have a wide variety of organizations individuals teams people that are all standing behind this and not having it controlled by a single organization and that is is it was
a hard lesson that i've personally had to learn um and i've learned through mistakes and moving forward i am never going to support a a company that i am either part of or running to basically own and control and hold hostage an open source project so everything that we're doing from an open source perspective we are not only putting it back out into the community we don't even want to control it we don't want to own it we don't want to hold it hostage we want that community to thrive we want other organizations other individuals and people to have have uh control over that um
that group that project and we believe in that wholeheartedly that every open source project at this point that we have released and we have been part of we don't control uh and we've done that and we're going to talk a little bit about that i think more as this this webinar continues but i think that that's that's very very important and to be blunt i challenge other companies to do the same if you love open source set it free i love it yeah i think it's an important topic to talk about because there are a lot of open source projects that were fantastic that have
gone away and it's been yeah i mean you have to pivot away from that product or that set of tools and it can become very challenging for the enterprise or even hbc to go do that yeah the only thing i'd add is that as a as a you know a product manager as a person that's concerned about solving problems for customers i find it to be one of the best ways for us to discover new use cases new ways to solve different problems because people are so unique and so creative that they come up and and try and do things themselves for me it's a
great way to stay in touch with exactly what our customers as ciq want at the same time as not holding it back and and controlling it now we can go shape it and do what we think we can go and do with our resources and our customers so it to me it's it's the ultimate way to understand the community at large and to serve them even better and robert you've you've talked about um you know every startup and every company to some extent is always thinking about product market fit and the way to get that product market fit if you have an open source project
again is to give that open source project to the community and the community will shape and and define what that project means and solves for them and you get instant product market fit when you have adoption of an open source project that a particular company doesn't own so it's it's so valuable in so many different ways i totally agree i appreciate that perspective guys so back on ciq how does ciq enable that traditional hpc stack i know that there are several different pieces of that but what does ciq provide in that so we provide so over over my career personally um i've had a fantastic
opportunity to create various capabilities that have been widely utilized in high performance computing ciq is basically pulling the innovations and capabilities that i've created among other people much smarter than i that have basically been able to add value to hpc centers to users to people that are building these systems commercial hpc and we're now pulling all of those together to create a stack a traditional high performance computing stack now three of the main pieces that we have been very close to in terms of innovation has been provisioning it has been the op base operating system that's running the resource and it's been containers and if
you take those three kind of major pieces you have the bulk of a high performance computing system there's obviously there's more components to this like a batch scheduler file systems and whatnot but this gives us the ability to start creating a turnkey solution for organizations from bare metal going all the way up through the stack and and that's where we're leveraging so i i'd really like to you know take a moment and kind of um go through if you don't mind zane and talk about some of these three kind of major areas that we're contributing to that high performance computing stack so to kind of
jump in if you don't mind absolutely provisioning is a really big piece of this and to kind of describe what provisioning is it is if somebody wants to build an hpc cluster or even any kind of cluster when you're talking about a cluster it's a whole group of computer systems and these computer systems are going to be built in such a way that they're going to work together to collaborately solve a problem and a software problem now if you just have a cluster that's fairly small let's say in the tens of nodes um you know it might be possible just to go through and handle
them like you handle any other server right whether you're using config management or you're installing them by hand and then managing them by hand but very quickly when we talk about these clusters we actually are scaling way beyond the tents and we're scaling into hundreds if not thousands of compute resources and when we do that all of a sudden the traditional kind of way that you think about maintaining a server needs to be needs to be pivoted a little bit um because it's just too many systems to do in this kind of traditional sense so we've been doing provisioning of operating systems and resources and
management of that cluster through a variety of tools i created one called werewolf back in 2001 and um werewolf brings some very unique uh features uh to the ecosystem and one of them uh probably the biggest one that it that it really changed and pivoted was the ability to have stateless compute images uh or nodes so what this means is let's say you have a rack of a hundred computers you roll this rack up into the data center you give it power you hook it up into the network and you literally can just start turning on nodes you don't have to actually install them you
don't have to configure any of them and you can just turn them all on just power up the whole rack and we've been able to build high performance computing systems from the software perspective from the moment the hardware is all set up hundreds of nodes and i'd even go as far as to say thousands of nodes we can actually have the whole operating system built and ready to run across thousands of computers in an hour or two and the whole thing is ready now for production and that sort of ease of use and scalability put together is really what was some of the defining characteristics
and problems that we were trying to solve initially with werewolf back in 2001 werewolf has done that very very well and as a result it is still used today and i'm still part of the project today and now we've got a number of different organizations that are also part of the project and including openhpc which if you're familiar in high performance computing that is a linux foundation project to to really help organizations create turnkey solutions for for traditional hpc so uh werewolf is the is the provisioning subsystem for this now one of the ways that werewolf worked and i don't want to dive too deep
in here as you can probably imagine i i i tend to blab a lot but one of the things that werewolf one of the ways that werewolf solves this is it uses a a node image and we used to call this the virtual node file system imagine a charute sitting in a directory on a control node somewhere and that becomes the representation for every node that's going to boot and that's how we've basically done werewolf for literally over 20 years now with werewolf version 4 which is fairly new uh the main pivot has been instead of using this to root this virtual node file system
um let's instead let's pivot to using containers and let's leverage the container ecosystem to be able to blast a particular container out to thousands of nodes and run that container directly on bare metal and that's one of the major areas in which werewolf um has just been so incredibly helpful and useful to to everyone in in doing computing uh the last numbers i heard we had like 50 of the top 500 um systems were running werewolf so a huge huge number in terms of how many people out there are running werewolf so there's a lot of systems out there running werewolf and imagine if you
can use the exact same tooling that you're using currently for whether they're user application containers or whether they're enterprise focused containers you have a ci cd pipeline you have all of your you know your devsecops um processes in place to validate those images uh imagine if you can now just import those into werewolf and then blast it out to you know thousands of compute nodes or tens of compute nodes it it makes it that easy and utilizing things like pre-configured containers just like we use containers today we can have containers a container for open hpc we can have containers for slurm lsf grid engine pbs
just for the scheduling side we can include infiniband we can include gpus luster support fpgas everything into these images and then we can have these custom images that are targeted to do different things people can now just import those images just like you do with a docker image import that image into werewolf and then say we want to leverage we want to build a slurm cluster that's running lustre um or gpfs and we needed to do x y and z yeah that's fantastic done yeah i i want to also point out i mean you've said it a couple of times but starting in 2001 that's
21 years of experience with this thing running in different environments and i think that's why it's grown so much and it's exciting that it's continuing to grow and gain functionality and popularity every time i get on the phone it seems like i'm talking to people who are using this today so it's exciting uh i know you're just jabbing at my age right now i i got it i got it i know so i think the main the main takeaways from that is i mean you're talking about it's it's stateless it's centralized it's flexible it's simple to integrate and then there's a lot of pre-configuration you
can do and add it to that cicd pipeline so i think that's fantastic i mean from what i've been able to tell it's also really easy to get it up and running right yep yep robert you came off mute do you want to jump in yeah so that's actually you know a lot of topic of discussion is where we're going to take werewolf and what's next and obviously the community drives a lot of that there's a lot of great folks in the community you know representing a lot of large organizations um what i would say is obviously a very simple gui type of graphical interface
to help people that are newer at these solutions to get them up and running extremely efficiently and effectively this is an extremely lightweight solution so it's it's designed specifically for this potentially adding some state fault capability to it um and some apis to make it you know more extensible and then lastly i would just say that you know potentially utilizing this platform as a large-scale imaging as a service solution for uh customers of you know rocky linux for customers of hpc so beyond just hpc we think we could utilize this lightweight capability at scale that's absolutely sorry zane i know i talk a lot um
that's a really good point robert as we've been talking to enterprises it seems as though that you can kind of generalize how people manage their resources based on size of the organization so we have on one side we have people that are doing physical installs doing system administration almost by hand on number of systems then we've got this middle ground where the majority of organizations seem to sit which is they're using some sort of kick start and configuration management based solution to manage that infrastructure at a larger scale and to be able to do um you know management of the config management system using even
something like git so all the configs everything is in this in git which is then being propagated via configuration management and you can manage a large number of resources but what i didn't realize until fairly recently was the other extreme right again at one stream is everything by hand going all the way past configuration management and looking at it at this extreme what we're seeing is organizations very large clouds some hyperscalers they don't do configuration management even well they do to some extent but for the most part they're doing imaging so every system is being blasted out as an image and what we've been doing
in high performance computing is extraordinarily similar to what's been happening at the extreme scale of cloud and enterprise where we're handling and managing this from an imaging perspective and and provisioning out those images of operating systems out to those those resources whatever those resources may be or require in an extraordinarily scalable and efficient means and that's that's one of the areas that i think it was really i just learned about this fairly recently that that's how they're doing it and i thought it was super cool how similar we've been doing it from an hbc perspective as well as the extreme scale so one of the
points that robert just made is what if we can change that barrier of entry using something like werewolf to help organizations even do more uh imaging as a service across their organization and maybe even couple that with configuration management to where we're actually we're supporting um even smaller enterprises to mid-size and large enterprises all using these same sort of methods and models so that's something that that yeah thank you for bringing that up robert that was a really good point yeah i think it's really interesting and we've been talking about this from an enterprise perspective for many years of getting to being able to as
soon as it's time for something new tear it down build new and build it up and i mean you got the blue green deployments and it's all fine and good in theory but rarely does it ever actually work the way you want it to so it's exciting to me to see something that there are people making progress down that path i mean i feel like this has been 10 plus years in the enterprise we talked about trying to get to that level we just don't see a lot of people making that kind of progress towards it so i'm excited to see where this goes as
well uh so i know and from an hp hpc perspective we have werewolf to do provisioning but you've got to provision something so that next layer up is rocky linux so let's talk about rocky for a minute so um i'm going to start it from the perspective of the relationship between ciq and and rocky uh and just to just to hit the elephant in the room um i'm the founder of of rocky i'm also you know founder and ceo of ciq so one of the things that one of the decisions that i made again looking back at our open source stance our goal is not
to own control or hold any of our open source capabilities hostage so we didn't want to put this in ciq we wanted this to be a separate project and a separate organization and and that's why uh i created the rocky enterprise software foundation but ciq has been along with rocky and with resf as a supporting um as a as a founding supporter and partner of the organization so right from day one ciq uh and myself and others at the company have been backing what we've been doing with rocky uh and and our goal again not to do anything in it from a from a corporate
agenda perspective we really just want to facilitate wherever this this this community wants to go our goal is just if anybody needs help we want to earn that business we want to prove that business and that's our that's our model and that's it so that's how that's the relationship between the two and that really leads to what we could be offering in terms of the operating system portion of this stack so we add some various benefits uh on top of the base rocky linux and the other thing i'll mention is centos is about you know in our approximation from everything all the metrics that we've
seen it's about 20 to 25 percent of all enterprise resources cloud instances containers worldwide so that's a big number it's a really big number now that has just been end of life and and rocky is stepping up to become that replacement to become that alternative to what centos has been and how centos has been trusted for so many years and we're hoping to do even to do it better in the sense that we're not there's not going to be a single company that controls that we're going to be rocky is going to be completely neutral from that perspective and we're going to have lots of
companies being part of that but ciq being one of those companies that's that's helping and sponsoring and partnering to bring value to customers we're there to help we're there to make sure that if anybody has any questions if anybody has any problems we can we can offer that we can become that enterprise level production support arm as well as offering some services now we already talked a little bit about imaging as a service but there's also something else that we've had a lot of really great feedback from with the community and customers which is repositories repos as a service imagine if a particular organization or
site can be able to easily overlay their custom changes to the operating system and make that operating system theirs for their organization all their customizations everything that they need for that organizations even configurations being pushed out default configurations being pushed out and managed via a a repo and that repo has a bunch of code around it that allows organizations to easily add packages into this easily build packages through that do cryptographic signing validation and everything that they need as part of that repo you can think of it almost like what satellite provides but there's no licensing per node so it basically gives you the ability
to manage to update and control what an organization's linux needs are without ever again without ever having to worry about licensing costs or anything along those lines so repo as a service we're getting a huge amount of positive feedback we're also focusing a lot on fips and other security accreditations that has really been huge for federal but what surprised me was there's a lot of companies out there that really prefer to know that those uh accreditations have been met and and that they can trust the trust the operating system so we are moving forward with fips we have a timeline if anyone's interested in that
timeline please let us know and we also have the ability to block security x exploits that use buffer overflows as the vector of attack and imagine if all past present and future exploits again using buffer overflows were mitigated we have that ability to bring that level of security directly into not only the base operating system but also containers containers running rocky as that base image so what we're doing from the high performance computing perspective is to add all of this into that traditional hpc stack and you can imagine that coupled with werewolf and then coupled about coupled with containers creates an amazing base platform for
doing traditional hpc yeah so saying i'll do the opposite this time so usually i take greg back from the open source to ciq and solutions we're doing i'm gonna go the other way since he told everybody about our repo as a service solution so um the other thing i would highly recommend is there is a special interest group in the rocky community that is specific for high performance computing uh every engineer at our company has deep understanding of high performance computing there's a tons of people in the rocky enterprise software foundation that have deep high performance computing background rocky's being tested on all sorts of
high performance computing apps and use cases and storage models etc i can tell you that folks in the testing community are high performance computing experts so what i would highly suggest is getting into that community aspects and really you know looking around and seeing all the resources there because i think you'll be massively surprised by the compatibility of rocky linux when it comes to high performance computing applications and solutions that had a totality you know i i completely forgot the point that i was getting to when i was talking about centos being such a big piece of the market at 20 20 20 to 25
percent of all enterprise resources in hpc it's closer to 70 to 75 percent so rocky not only has a lot of interest in enterprise hyperscale and cloud but the interest in high performance computing is massive and if you look at who are the people that are part of that rocky community two of the people for example the two people that lead the testing group are both from hpc centers so we definitely have a large amount of uh interest and background with the high performance computing section of the industry and again just to reiterate what robert just said we have a sig hpc a special interest
group specifically designed to further enable high performance computing on rocky and if this is something that people are interested in we will post the link to the uh to our our matter most which is basically our open source slack that you can join in and then you can search for the channel that's sig hpc and please join uh we'd love to have you as being part of the the rocky community excellent thanks greg so now we have provisioning we have an operating system now it's time for people to start deploying apps and i mean nowadays everybody wants to talk about containers and ci cds so
i know you kind of alluded to a little bit earlier without tainter but tell me about obtainer where does it fit in this model and how does from the we've got the base we've got the os and now what so apptaner started life as singularity and a lot of people within the high performance computing world know singularity very very well it it has been the absolute dominant market share leader of containers within hpc if you're not in hpc it wouldn't surprise me if you've never heard of singularity but if you are part of the hpc community i'd almost put money on that everybody has has
heard of this and and potentially even been using it because it has become such a prevalent and critical piece of high performance computing infrastructure and this is this is because when containers were really starting to evolve and and gain traction and publicity and interest in the enterprise and cloud space the the the methods for dealing with with the run time for containers were very root kind of centric and it didn't allow very well for a beowulf style architecture that may have hundreds or thousands of users logged into them with none of them having any privilege or access above what they would have as just a
generic non-privileged user to give them access to a container subsystem that was all running as root uh would have would have defied all of our security aspects and we would we would basically you know that would be a security exploit if a user has the ability to control the host operating system and that's exactly what docker would have given them at that stage so that plus several other facets you know in terms of how do we make good use of gpus infiniband how do we do things like mpi which is the message passing interface which is the library in high performance computing that we use
to enable parallelization of compute jobs across node boundaries so we can actually have 100 nodes all running the same application and doing you know uh um communication between those nodes in a very highly optimal manner on whatever network backplane you have um we had to figure out how to do all of that how do we integrate with a batch scheduling system that is controlling all the resources of that hpc system so long story short we really needed to come up with a way of integrating containers into an hpc system that fit with the hpc architecture and use cases and um and that was what singularity
did as a result of that within just a matter of months singularity went from just an idea to it was installed probably on the majority of of big hpc systems that supports containers again in in a matter of months not years so uptake was incredibly fast it solved a critical problem it was highly optimized for high performance computing use cases and architectures supports things like mpi infiniband gpus file systems any batch scheduler that you want allowed non-privileged regular every or average ordinary user out there to be able to leverage any container they want and so it really took off and we've done things even differently
than the traditional kind of container ecosystem like we have single file based containers that we can cryptographically sign and encrypt so we actually bring even a level of security that differentiates the container supply chain further than anything else to my knowledge that's even out there today so singularity made a bunch of great strides we really wanted to ensure that singularity stays open source and in the community uh for the long haul and we decided in a collaboration with the linux foundation to bring uh singularity into the linux foundation and allow the linux foundation to take over ownership and control of the project and their one
request that we had was well that's great everybody was on the same page but they just asked us to rename the project and so they can properly protect the copyrights and and trademark so we renamed the project we went out to the community and we said uh we need to rename the project if we're going to move this into the linux foundation and we gave a few suggestions a few other suggestions were offered and apptaner was the voted lead so we ended up renaming the project to optaner so singularity at this point has been renamed to aptaner we have a thriving community around this we're
part of the linux foundation and it is the goal of this and the goal one of the goals of moving to the linux foundation was how do we better cross-pollinate between the high-performance computing world and the enterprise and cloud world in terms of container and container infrastructure and so we've been focusing a lot on um that cross-pollination and leveraging those those resources on both sides excellent so robert i know you're very passionate about the security aspect of apptaner singularity and kind of that single file being able to have the cryptographic signature on all that and you want to dive into that a little more yeah
i still remember the first time greg explained to me in detail what he'd created i said wait let me get this straight greg so basically you're saying that i can take my application and move it to any cloud any linux operating system and running on any piece of hardware or anywhere i want in the world and not be completely locked into any cloud or anywhere or any specific piece of hardware ever again and he said yeah and then i said wait a second so you're also saying that it's immutable and it's perfect for every kind of research that's possibly available out there yes and it's
immutable oh okay and reproducible oh okay you literally just solved like 25 problems that i've heard from customers over and over and over again right so it's kind of like i heard everybody talking about serverless and i'm like why would i want serverless when i can just move my apps anywhere i want in the world why would i want to be locked into anything right so this to me is huge greg said something that's really important it's software supply chain security aspects we can cryptographically sign every step of that process and we can also sign for security signatures for you know deployment signatures devops signatures
etc and then with our new platform we can actually even ensure that that doesn't get run anywhere unless it has all the correct signatures so um so this was a basis of a lot of the future steps that we took and it's all because of those 25 problems greg solved with that one as greg says pandora's box that he opened um and that led us into you know creating the next generation of high performance computing that we've been working on as well and we'll talk about that next week though excellent so how does ciq bring all of this together what does it actually mean for
a customer some feedback that we've gotten is and this is both from vendors as well as customers they want a turnkey solution they want to know that they can come to a single location and get the the entire foundation of what they want to build and and know that that can be supported and it integrates easily so we looked at all these different pieces that we've we've created and we're continuously maintaining and and contributing to we said okay let's put these all together into a turnkey into that turnkey solution and let's build a traditional hpc stack specifically to be easy to deploy highly scalable very
very flexible and become that base foundation for hpc centers hbc consumers and even enterprises that are looking at doing compute focused workflows and let's simplify that let's democratize that in such a way that anybody can go and and do this and what's what i love about how we approached this was we were not going to make any sort of architecture uh statements or or compatibility options i i don't know if somebody wants to use an nvidia gpu or an amd gpu i don't know if they want to use an fpga i don't know if they want to use slurm if they want to use lsf
or grid engine or pbs i've got no idea what a particular customer is going to want to use what i want to do is i want to provide them that infrastructure substrate that they can go and build the tools that they need to create and solve their problems and that's what we've created we've created the a turnkey solution that can go and build up an open hpc cluster or a custom cluster or however they want to build it easily and and efficiently and scalably so it's very simple it's flexible it's capable it's scalable it's secure and that was our goal to bring all of that
together uh we we do have some organizations that are uh leveraging us for kind of an a la carte stack in a matter of speaking some customers are strictly for singularity and apptaner other customers are focused on werewolf and or rocky actually we have a lot of customers on rocky right now so it could be a la carte as well however an organization wants to deal with this we're here to support the organization and and i i it's one of the things this is a little bit off of topic but i think it's really super important to mention uh you know we are offering both
capabilities and licensed values on top of open source pieces as well as offering services and support and for services and support most organizations are pretty much doing what what the legacy model has been and when we first started thinking about how would we do this we decided to completely rethink and and be willing to throw out every idea that's not working and that's not beneficial to the customer and what we decided to do is make sure that everything we are doing is about enablement of the customer and if it's not enabling the customer if it is a value to us and not the customer then
that's not a win for us so i can give you an absolute easy example of this most or many organizations will charge for support based on core or socket or even nodes and some organizations really have to use nodes to get it through their procurement department but most people don't like it they don't want to have to count the number of systems or resources that they are paying for support they don't want to do that so who does that model help it helps the company who's offering the the support not the customer so we basically you know if a customer really came to us and
demanded per node pricing we'll do it but we want to make sure that we were offering something better even site licenses can be difficult so we decided to completely rethink this and here's what i can say every person every organization and team that we have shared this model with even the teams that said we never want support because we have a great team and we don't need the help and we don't need the insurance policy even those companies are now joining with us because they love the model and they see how we can now help and add value to their team as people as individuals
and so if anyone is curious on exactly what that looks like reach out to us and we'd be happy to walk you through it absolutely so i mean guys if you have any questions go ahead and start posting those i know robert you were you were about to say something yeah i i take it one step further so you know greg told you about some of our design principles you know one of them was being modular um in how we approach this and you know having the building blocks but also making it all work better together as well over time so if you were a
dell customer of dell hardware you could utilize their omnia stack and utilize werewolf rocky and apptaner if you're an hp customer you could potentially utilize rocky then apptaner and maybe keep some of their you know secret sauce that they utilize if you're a customer of penguin you can utilize rocky you can utilize apptaner in a traditional stack like many folks do today and that's just an example of some folks you could partner with super micro you could do the exact same thing you could use our whole stack or you can use pieces of it that makes sense to facilitate the additions to what you're doing
so definitely work with your oems one of the overriding principles that you know greg and i talked about a lot about all of our open source projects and all of our projects in general was to have every cloud every oem every component of a server and every isv that could potentially run on any of those environments as partners as sponsors as folks that we work with on a regular basis and it's not by name if there is an issue within with a system i want to be able to call that oem from our support group explain to them what's going on where we want to
not uh make it so that the customer isn't getting served we want people to get their problem solved and if there's a driver that needs to be executed on or something that needs to be done we're there to help them do it and then the last thing i'd say on containers obviously there's a lot of libraries and os capabilities that get built inside of there we have a full staff and team to enable the entire stack that we're talking about and giving your application speed and life and and flexibility at the same time as giving your end users and your admins complete control so that's
that's what we were doing when we were designing all of our software um and potentially in the uh in our open source solutions as well perfect well guys there aren't any questions i know i just heard greg's reminder for his next meeting pop-up i know you're a busy guy so i want to respect your time and i appreciate you joining uh again we ran out of time and didn't get a greg's story sorry maybe next time we would push for that or maybe we could just force robert to tell a story we'll see we'll see where we get anyway appreciate you guys joining again go
like and subscribe and and we'll see you probably next week i guarantee there'll be a story next week awesome thanks guys bye everybody you
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.