Optimizing Linux for AI/ML Workloads
CIQ's product team gathers for a live tech session on RLC Pro AI, the Rocky Linux from CIQ variant being built for AI and machine learning workloads. Host Rose Stein is joined by Brian Dawson and Brady Dibble from product, Rob from strategic projects, and AI advisor Zach, who together explain why a general-purpose enterprise Linux leaves expensive accelerator hardware underused and why production inference is where the cost shows up.
The conversation covers what sets RLC Pro AI apart from base Rocky Linux: a foundation on the upstream 6.12 long-term kernel with the option of tracking stable releases, build flags tuned for AI rather than general use, out-of-tree drivers and hardware enablement for GPUs and other accelerators, and frameworks such as PyTorch and TensorFlow baked into a reproducible gold image. The user space stays compatible with Rocky Linux from CIQ 9, so existing software stacks continue to work.
Infrastructure teams, Linux administrators and AI engineers running or planning production inference walk away with a clear picture of the trade-offs involved, why memory bandwidth and I/O matter for tokens per second, how confidential computing fits multi-tenant deployments, and how to join the tech preview program ahead of general availability.
Key takeaways
- Roughly 94 percent of AI workloads run on Linux, yet most run on general-purpose distributions that are not optimized for AI.
- RLC Pro AI starts from the upstream 6.12 long-term kernel, with stable releases as an option, to expose new hardware support and performance gains quickly.
- Kernel build flags tuned for AI rather than general-purpose use can have a large impact on performance at scale.
- Frameworks such as PyTorch, TensorFlow and Hugging Face libraries are baked into a reproducible gold image to cut setup and dependency work.
- The user space remains Rocky Linux from CIQ 9, so existing software stacks stay compatible while the kernel moves forward.
- Tech preview signups are open now, with a preview image planned ahead of a general availability release later in the quarter.
Questions this video answers
What is RLC Pro AI and how is it different from Rocky Linux?
RLC Pro AI is Rocky Linux from CIQ built specifically for AI and ML workloads, primarily production inference and tuning. It pairs an upstream 6.12 long-term kernel, AI-specific build flags, out-of-tree drivers and pre-installed frameworks like PyTorch and TensorFlow with the standard Rocky Linux from CIQ 9 user space, trading general-purpose features for performance.
Why does the Linux kernel matter for AI inference performance?
The team explains that memory bandwidth and I/O often limit tokens per second during inference, and a generic kernel may not fully use the accelerator hardware underneath it. Newer upstream kernels bring driver support and performance improvements that would otherwise take months of backporting, so tracking them closely keeps expensive GPUs from sitting underused.
How can I try RLC Pro AI before it is generally available?
CIQ is accepting signups for a tech preview program while RLC Pro AI is in active development. Participants get the bits to deploy and test, with direct access to CIQ engineers for feedback, ahead of a preview image with most planned features and a later GA release. Signups run through ciq.com.
About this video
Recorded on May 23, 2025. CIQ had recently introduced Rocky Linux from CIQ AI (RLC Pro AI), an operating system optimized explicitly for AI/ML workloads. In this tech session the team walks through how the kernel and user space were engineered to get AI projects running faster and more efficiently.
The session covers:
Kernel and User Space Enhancements How RLC Pro AI enhances the kernel with AI explicit build configurations, support for Confidential Computing, and support for hardware acceleration components.
Key User Space Additions How RLC Pro AI includes both proprietary and out-of-tree modules alongside up-to-date drivers, eliminating post-installation headaches.
Performance Optimizations How the RLC Pro AI kernel is configured for AI-specific performance optimizations inspired by high-performance computing but tailored for AI/ML workloads.
This video is part of the RLC Pro AI playlist. Browse every CIQ video by product and topic.
Transcript
Good morning, good afternoon, and good evening wherever you are. Thank you for joining. At CIQ, we're focused on powering the next generation of software infrastructure leveraging the capabilities of cloud, hypers scale, and HPC. From research to the enterprise, our customers rely on us for the ultimate rocky Linux, Warewulf, and Apptainer support escalation. We provide deep development capabilities and solutions, all delivered in the collaborative spirit of open source. Okay. Hello and welcome. Oh my goodness. Look at this team. This is very exciting. Hello everybody. Thank you so much for being here. I'm really excited to talk about RLC AI, optimizing Linux for AI and ML workloads. But first, let's say hello to everybody.
So, you guys probably know me, Rose Stein here, working at CIQ in the sales department. Very glad to be here with Brian Dawson. Unmute and introduce yourself. Unmute. Yep. You got to unmute. Unmute. Okay. Maybe we'll come back to Brian. Brady, you're up. Hi there. Brady Dibble, product CIQ. That's it. And I'll come in. Short and sweet, man. I like your style. Okay, Brian, how's it going? I am Brian Dawson, product at CIQ and uh I ran into Chrome Tab Tabitis and I lost my Streamyard window, so I am now here. Thanks for having me, Rose. Chrome Tabitis. You heard it here first. It has been coined by Brian Dawson.
That is awesome. Thank you so much. We all know exactly what it is that you're talking about. Where am I? Where's the voice coming from? What's happening? That's awesome. Well, very glad that you are here. And Rob to follow. How's it going, man? Uh, doing well. Rose, how's it been? Been a little bit. It has been a little bit. I I see you've got uh we've got the same letters, but different colors. I'm liking. Yeah, I I got to get some of the new gear, I guess. But no, I think this is this I got this with the last one. But yeah, so now my little title says head of strategic projects.
Here I am again doing new things. There you go. Yes. That's actually one of the fun and wonderful things of a startup is that there's movement, right? You kind of you come in, you're like, "All right, this is what I'm doing." Just kidding. This is what I'm doing. I I think I've had about a half a dozen titles since I I started here three years ago. But yeah. Yes. This is my favorite one yet. Okay. Okay. Maybe we'll just kind of like settle in here a little bit. I like it. Well, glad that you were here. Thanks, Rob. And of course, Zach, so you're actually um I think the newest addition to CIQ.
View full transcriptHide full transcript
Uh so tell us about yourself. Yeah. Hey uh so I've been advising the team here on AI applications for for Rocky Linux and many other projects. Um I am the uh head of AI at AI Insight Solutions which is a consulting firm but I've been working very closely with these guys. I'm quite excited to see what they've built. Yes. So this is a hot topic you guys because AI like what is AI? There's lots of feelings and emotions around AI, you know, and then there's the technical details. So, we're definitely going to dive dive a bit more in here. So, um I think who's going to do that like just that RLC AI overview?
Kind of tell us what what are we talking about here? Well, I I'll start and then I'd love for Brady to chime in. So, you know, I think we're here because last week we announced that we are actively building Rocky Linux from CIQ for AI. Uh and what that is is well we realize first that uh I think it's 94% of AI workloads are running on Linux systems. However, a good majority of those are not at all optimized to support AI. They are generalpurpose Linux distributions and in fact many of them are not even enterprise Linux distributions with sort of the enter enterprisegrade um um upstream support.
Um we also realized that um inferencing is very expensive. Um and those two things uh present an opportunity for us to build and deliver RLC Pro AI to help people running uh or delivering AI workloads at scale via production inferencing um to reduce cost, optimize performance and reduce the overhead in customizing, modifying and deploying Linux to support those AI workloads. and Brady and team here have identified a number of features that meet that target or meet that need for uh for those people running AI at scale. That was a lot of words Brady. What did he say? So the past couple of years AI has exploded and everybody has to have an AI project and a lot of the focus and discussion has around AI infrastructure has been about getting from zero to one.
CIQ works with a number of customers who are they're already in production and they get it out into production, they go, "Wait a minute, I'm not going as fast as I want. Wait a minute. I've got all this expensive hardware. Why isn't it going as fast as it set on the brochure? Uh, my Ferrari is, you know, stuck in first gear. Um, or wait a minute, I have to keep bringing down like new CUDA drivers all the time. I have to keep dealing with dependency, you know, heck with all of these new like drivers all the time. compatibility issues every time I update. I have to rebuild this OS constantly to keep it fine-tuned.
And we found that we could actually support a lot of our customers with making that just go faster with the RLC AI. There's a like three four main areas that we're focusing on. The first was all of the nice new stuff is in the upstream kernel. So the first thing we were doing is bringing down the new kernel to say, "Hey, everything you want is already in that upstream kernel. We're going to make it easy to access everything as soon as it comes out. The second one is your generic kernel, the off-the-shelf kernel, that's for general purpose. So even things like changing the build flags that are used when you're building the kernel can have monumental impact on your performance at scale.
And then you know we're start bringing in there's out of tree drivers and there's other optimizations, there's frameworks, there's compatibility that you have to deal with. All of that goes away with RLC Pro AI. Wow, that's a bold statement. All of that goes away. Zach, do you agree? Yeah, I mean that's that's the whole point of it. Uh I mean, as a consultant, I spend a huge amount of my time dealing with dependency hack. um installing drivers, just getting a fully compatible stack from your drivers to CUDA uh to the CUDA toolkits like CUDA DNN to PyTorch to Python like building out all of that yourself every time is difficult enough.
Keeping it up to date um is just way more work than I ever want to do. Um, I love systems where vendors such as all of you, uh, all of us package that up and you can just install it and and run with it. Um, there's it's it's work nobody wants to do. It's work that's great to do once and then push out to everyone. Nobody wants to do it. Huh? Is that true, Rob? Yeah. and I'd add that and it requires a special expertise that frankly a lot of organizations that need to move quickly don't necessarily have. Right. Um right. So I think a a lot of you get a lot of people honestly CIS admins Linux CIS admins that have to do this to support this core business initiative.
Um and they don't know how. So yes, I don't want to be burdened or straddled with ensuring um uh you know with with owning the foundation for my company's AI workloads while I don't have that expertise. So hence CIQ hopes to uh empower those those people. Yeah. And and there's like also there's a a gap between on the one hand you've got some researchers saying oh I want to use flash attention to make this model faster but the CUDA version we have doesn't support it. And then you're going out to a system and speaking a bunch of words they don't understand because they haven't heard of flash attention before.
And and there's this whole compatibility layer of there's some feature in the newest version of torch you want to use, but your stack hasn't been built to support that all the way through on the hardware you're on. And so the feature you want doesn't work. But you're out here at this end, you know, in PyTorch land and your system is over there in hardware land. And there's a lot of steps to cross in between to to get to the point where you're speaking the same language and you're you're you know on a stack that's doing what you want in user land. User land. So I mean there's got to be a reason right?
There's got to be a reason that um people you know not a lot of vendors are doing this. What what what is that reason? I'm gonna pick on Rob. Right. Rob, you I think this is more of a Zach one to be honest because like you you've been all over the place, you know, especially inside of this company and in other companies. So I think I don't know. Okay, you want to pass on to Z? No, I mean I think that when you really look at it from my angle. So where Zach comes down from like the top down, I come from the bottom up, right?
So exactly. So when you look at the the hardware and a lot of what's going on there, one of the biggest challenges in deploying AI right now is the hardware layer that's out there. How fast it's growing and those challenges around it, right? So you have these big production environments where there's a lack of timely support for Linux kernels. Many of the drivers, the MPUs and other accelerators that are coming out um are completely dependent on what's going on with Nvidia, AMD, Google, ASIC, and all these other ASC kind of providers, right? So you have kind of this lag behind the mainline uh releases that require like custom patches.
So when one of these people come out with this new tool that just landed on the market as Zach kind of talks about you now have to figure out what that whole dependency tree looks like all the custom patches forcing teams to choose between different kernels stability security and then just the hardware enablement in general. So not all kernels are created equally um especially when it comes to like the IO performance and with when the workloads are running on those. So while you might have some really great hardware out there, if the kernel isn't using that IO, well, you could be looking at bottlenecks within that area um and how efficiently you can move data, right?
The core of all this is data, data, data. What's that need? Memory, hardware, and the ability to push between those. So efficiently moving that data between storage and the accelerators and the drives, all that's a really big deal. So that's kind of what the core of what we've been digging into for the last, you know, four or five months here is just what that looks like and how can we better accelerate that for customers. Yeah. And there are vendors trying to solve this problem, but a lot of it is, hey, here's here's a Docker container. Um there's not that last step down to the hardware level down to the to the kernel level.
Right. I I would Rose just to chime in. Um look there this we are move so technology adoption and penetration over the past couple of decades has continually increased from looking like this to looking like this. We have this wellfunded AI race where we have multiple layers of a complex problem for companies to figure out get deployed to scale. I remember I was watching a talk at a machine learning symposium recently and um uh God it wasn't Amazon I am losing it right now let's say say Zixedia and she's telling the story of she presented uh research work to uh the company execs fully expecting she'd say she needed you know two more years to develop this and productionize it.
They told her they wanted it productionized in a month, right? They went ahead and did it. It saw success. But you have this complex stack of problems that people have to deal with. And not surprisingly, people are ignoring what's happening down where where the software touches this valuable hardware, the OS layer, right? Um, and so we're here to solve that, to call attention to that gap and help close that gap out. That's awesome. So tell them a little bit more like what you guys have already touched on it, but I just I think this is a really important part of RLC AI is what is different from like the base Rocky Linux image.
Exactly. U that's a Brady. Let's Yeah. What's different? Well, when you're taking a a distribution broadly especially for enterprise you don't necessarily know what the customer's use case is so enterprise Linux distributions are designed to be somewhat general purpose and then the customer is expected to add on additional optimization themselves um or they have to go through professional services to have someone else do that maintenance for them and you create a custom one-off CIQ's like workload specialized products RLC Pro AI and the recently released RLC Pro Hardened that's doing a lot of that out of the box. So you have a gold image that already has that heavy lifting done for you and is reproducible in a way where you do not have to be the one constantly keeping up with all of these drivers, all the compatibility, the optimizations, any little challenges with build flags.
So the first difference is saying we're going to take this and just make this optimized for AI use cases out of the box. these specific workloads, tuning inference primarily because that's where you're you're not going to get the most out of your OS unless you're literally building like at build time specifically for this use case because there are you're turning off a lot of things. There are trade-offs essentially. You're not going to use this and then go run a regular web server. This is specifically for AI use cases and you're using it for tuning or using it for inference.
And when you are, as we talked about, when you're actually productionizing your AI workload, you finally get it out, you've done all of this work, you've set up your AI infrastructure, you've got it in production, and then you look at your bill and you go, "Oh my goodness, as we've probably heard, every time someone pings a, you know, an LLM, that's that's costing money." And so any any improvement performance, that's saving you money. And simultaneously if you can improve things like latency that's improving the customer experience that lag between when they you know submit a query and when they get that response back that's massive in the customer experience.
So making those um building from the beginning with the idea of this is a single-purpose use case to an extent of we're maximizing your AI infrastructure for AI LM and ML workloads. Just starting with that idea is already a major mind shift. Mhm. Then we've are working on now when we get into the technical specifics, that's when we start talking about specific build flags, specific feature support, specific hardware support, um the GPUs and all the other PUS that are being utilized, making sure that that's not just compatible, but you're getting the most bang for your buck. Um and performance, like performance is absolutely critical. But what now you're in production, now you've got other concerns.
Um like what happens if it breaks? How do I get support for this if I'm just pulling things from the community and doing it myself? Uh what about security and compliance? We've seen along with this push for performance also now that it's in production. Like I said, get it out in a month. Great. It's out in a month. Now make sure it's secure and compliant. That that suddenly that's a bigger deal. Um, encryption at runtime is like with confidential computing has become a hot topic this year because especially for like multi-tenency scenarios, if you're running in a multi-tenant environment, you need to make sure that that uh your work can't be seen by the other tenants.
So that's where things like confidential computing come in. Um, that is now Yeah. Now you're in production. Uh, you don't want security and compliance to be an afterthought. No, I'd call out and I maybe Rob is better to talk to it or want someone else here, but we identified um that the absence of the latest upstream kernel and some of the AI focused Linux distributions out there today is are really hobbling um um teams. So I think a key feature where we say what is a major difference between RLC Pro AI as we call it for short and Rocky Linux um um it's going to be how we handle the kernel updates and pulling the latest upstream kernel forward.
Yeah, that makes sense, Brian. And and to that like that's some of the things we really like that when we started digging in. Uh Linux is just now really starting to truly like embrace a lot of what's going on with this. like they're they're getting up there and they're looking at like the latest like 613 uh builds and such start looking at like some of the ARM support we need like ARM 9A where you really start to see the ability to uh address memory much faster. And so the core things we really looked at is like what is when we started digging into this is like performance benchmarking and validation around those kernels and then what can we do to be better about rapid adoption of new hardware as it comes out because it's going to be very important to us.
And then obviously the thing everyone's here for ease of model deployment when when using this. So those are the three core things we looked at when we really started digging into what we're going to do with this um kernel and this entire product. And and it's like one one of those things you said on like memory paging speed. Like that's especially inference time. One of the biggest factors limiting like throughput tokens per second coming out of your models is your memory bandwidth. Yes.
Like that's where it really pays to have the whole stack optimized for look I have a huge amount of data in RAM and I want to be able to access all of it very very very quickly and and that's not like that's not a problem you know that's a non-trivial problem to to solve definitely we and especially when we start looking at these bigger models where they start thrashing memory the ability to page out and such much faster is going to be very a big deal when you're talking about like warm versus cold startup of these these models So, one of the things that that you mentioned, I think Brian is just talking about like how many people are using Linux that are running AI workloads and it was all a vast majority.
But how many people are running AI that are that are at all? Like almost everybody. I don't know. Is there anyone who's Zach since my name was on the front of that I'm gonna say I don't have a specific number I probably do. I can't recall it. Um but I think we know I I will share one thing before I go to hand to you Zach. Um um we know venture money fuels a lot of this innovation right uh and people that aren't venturebacked um we already know uh you know the big fangs are all leaning into AI. I was reading an article yesterday that said over 50% of venture funding last year went into um AI focused companies.
Um so I will grossly say before I hand this out the bigger question is who is not uh engaging with and running AI workloads. All right. But I'm gonna go to the expert. I probably should let the expert go first. Yeah. I mean I don't have a statistic off the top of my head. Um I have anecdotes and as we all know the plural of anecdote is not data. Um but every every single client I consult with is thinking I mean I don't I don't think there is a company that isn't thinking about AI to some degree like it's going to affect the entire entire economy.
Um if my mom is running an AI app on her phone it's literally everybody. I I went in for my annual physical yesterday and my doctor had an like I asked him a question and he's like, "Hang on a sec, like let me check the research." And he's got an actually very cool app um for searching scientific articles and summarizing them for doctors. Um and I mean this, you know, this guy's an internist at MGH here in Boston. You know, I've never talked to him about AI before. I'd only talk to him about, you know, hey, keep keep working out and, you know, whatever. Um, and it's amazing.
And but what I will say is is you got to remember like for every person using AI, there is, you know, a company on the back end managing all of those servers that that run it. And so that's that's that's the thing to keep in mind here is it doesn't take many companies deploying AI to reach a huge number of users. Like the scale of these deployments is massive. And even if a given company isn't necessarily building and deploying their own AI systems, they're almost certainly using someone else's AI systems. I mean, it's getting baked into everything from, you know, Google Meet's transcription features, which I really love, and there's a million products doing that, um, to all sorts of other things.
So all all of those cases like as the scale grows exponentially the compute needs grow exponentially um and you know there's it it it takes a number of years to put those new data centers out in in the desert. It's a much better idea to get to get more out of your existing hardware. Yeah. The the thing I'd say on this um and talking with some hardware vendors that I know they they kind of say hey this this kind of goes in like 10ear bubbles and we're kind of in the 2020 bubble right now. And right now, if your technology company isn't looking at how they can either use AI or ML style loads, they're going to be behind the the eightball with a lot of what's going on.
Um, if you have any kind of like forward-looking strategy, you should always have something at least looking around in there and understanding how it can help uh be applied to your work your work or your uh your uh just work. Yeah. Well, so I mean that makes sense with using the upstream kernel because then you get the newest drivers and hardware and like the most newest stuff because it's changing so quickly, but like how upstream, right? Like there's lots of versions of upstream kernel is there a specific one that is being used and then how long does it stay on that one? Do you have to do like a whole, you know, reboot every time?
How often is it? I'm so curious about this. Yep, we are getting CIQ has a a strong relationship with the upstream and we are getting a bit forbit with what is kernel.org kernel the straight same kernel you would take from kernel.org That's what we're building and optimizing for because the pace of development is so quick that the moment you start forking that upstream, you are now to an extent behind. Uh we're see the RLC Pro AI is starting with the 612 long-term kernel as its foundation and we're also even looking at the stable as an option for those who really want to get access to things as soon as they drop.
Um, as Rob mentioned, even in the 613, the 614, some of these more recent stable builds, there are massive um, performance improvements that would directly impact not just like not just the bottom line and the cost of running AI workloads. It will also unlock features that you previously like prevented entire strategic initiatives. You can do things that were not possible. And it can take some of these require would require backporting massive components and subcomponents that may take months and months and months or even years. And in that time, the industry has moved moved on. If you know, if you're waiting even a week or a couple of months or six months to access the latest performance, latest features, you're you're behind the eightball.
Um, so getting um getting these newest features and newest drivers and newest performance enhancements directly to customers in a way that they're confident about running it in production environments at scale. And we're talking at scale. Uh that's one of the focuses of RLC Pro AI is bring all of that latest cutting edge technology um into your foundational OS so you're getting the most out of your cutting edge hardware and using uh that as a stable foundation for your cutting edge software stack. And if there's one group of people who is just relentless in their push for I want the latest and greatest thing. um whether it's an innovation for speed or an innovation for model accuracy, it's it's AI engineers like they are relentless.
You read a paper and you want to go implement it and you don't want to go start compiling your own kernels for CUDA to do it right or have to go beat up on your infrastructure platform ops or DevOps team, right? I did want to add I think you know a key thing as I started to dig in and research more to help define um the feature scope uh for RLC Pro AI. It became clear to me that while right now Open AI and ChatGP dominate the production um AI workloads, again this is something like over 80% of of of what we see of AI out in the wild ultimately is APIdriven through chat GPT that is rapidly swinging and for a couple of you know two reasons.
One is very similar to what we saw with cloud adoption. This is great. I can deploy fast and easy. I'm going to consume it. And then people stopped and said, "Oh, wow. Wait, this is really expensive. We need to be more thoughtful." Right? So that's causing people to to to kind of come back into their own data centers, into their own virtual private clouds and hosting their own models.
Um two um is as as AI becomes less of an experiment and it be and it empowers businesses more critical IP um uh uh uh and specialized use cases around critical IP are being deployed so people need sovereign AI on premise um and then third we all heard about deepseek um we've seen uh the rate at which um you know meta open source and llama drove adoption and now we have models open source models like Mistl etc that are rapidly involving evolving and now able to compete with some of the services you would have gotten from open a or anthropic so what this means is
um people have been able to ignore figuring out how do I run this on prem how do I manage my own AI workloads but they now have to rapidly figure it out yeah and it's important to recognize like in AI the open source isn't that far behind the closed source state-of-the-art. Like we all remember chat GPT's launch and what an amazing product that was the first time you used it. And like you can run a model that was better than the original chat GPT on your phone right now. And so just imagine going back in time and saying to yourself at the launch of Chat GPT, hey, you know, the next phone you buy, you'll be running Chat GPT on the phone's GPU.
Um, and that took some software innovation in the models, but it also took a lot of hardware innovation in the phones. I mean, that's the whole reason Apple moved to 16 gigs, right? Um, is because they want to be able to to get Apple intelligence rolling. So, I mean, you're seeing it across all the hardware. Yeah. But like, there's there's open source small models like I forget the app, but the you literally can just download the weights. It's a few hundred megabytes and run them on your phone. It's absolutely brilliant. Oh, yeah. So, Zach, can you define inference for me in this context and why it's important?
You guys have said that word a lot and I really should just on the side chat GPT it, but Yeah. Yeah. Well, actually, that was the example I was going to use. So, when you use chat GPT, you're you're doing inference. Um, so so there's a model running in OpenAI's data centers. You type some text into a box that gets packaged up into, you know, an API call that goes out to some server and then that server does a bunch of matrix algebra to decide what words to send to send back to you. So that's that's called inference. And inference is important because that's what powers every single AI application that's interesting to humans like you know my doctor or your mom or me or all of the AI use cases.
Even you know I'm an AI researcher and AI consultant and 95% of what I do is is inference. Um that's just models get trained once and then deployed for potentially years after the fact. And when you you amarize that, you end up using a lot more compute on inference because inference is what use is useful. Inference is the entire point of training a model. Um if you train a model and you're never going to infer from it, you know, maybe that was an interesting exercise. Maybe you got some papers out of that or something. But inference is fundamentally what makes models valuable. And if your belief is that AI is going to power more and more of the economy, that's a belief that there's going to be more and more inference.
I appreciate that. Thank you. It does coming together and making sense of why that's important for RLC AI as well. Cool. So, what else do we want to talk about, guys? I feel like we kind of covered a lot, but there's probably more more. I had I did want to extend on my other statement and Jesse and I'd also like to get Zach's thoughts on this, but look, we are making the bet. I talked about this move from from um uh API uh uh uh implementation of AI and AI inferencing to onrem. Um I I I do want to call out that and again similar to cloud workloads, right?
I think where we're going to see a steady state is is not all onrem self-hosted open source. It is not all um through a third party but rather uh at scale or we are going to deply deploy our AI workloads in a hybrid manner. uh and as I look at it from a product standpoint to me what that means is we have to offer equal priority to being able to deploy uh on bare metal and on-prem environments um support containers to be orchestrated by Kubernetes as well as delivering cloud images to efficiently support cloud workloads um and again I think I I think we are seeing the industry move towards this hybrid model I'd ask the rest of the team um if if they agree with I'll take that as a note.
Yeah. And I know I'm sure Zach has some comments, but absolutely we've seen that there are certain use cases that just don't make sense to build out an entire infrastructure. Um, and if you want to test quickly, the fastest way is to go through the cloud. The cloud has the hardware, they have the stability, they have the processes. If you are trying to spin up a an operation, a production inference model in a very short period of time, you don't want to wait for a new data center to be built or provisioned. Uh you can get that up in the cloud right away. And you will always be, you know, using the cloud at some level.
Now, once you've really got an established process, maybe for cost savings, you're looking at longterm to move things on prem. Uh but when you're when nothing beats the cloud for speed uh to deployment. Yeah. I mean, right now, um, Microsoft and and Meta are two of the biggest companies buying up the most chips right now. And where is all that going to go? Cloud. Yeah. And you can't if especially if you can't get your hands on some of the newest technology, some of is not only very expensive, um, but there also have been supply issues. If you really want to run on the truest cutting edge, the latest blackwells, uh it makes a lot more sense to just go and utilize what the cloud has available.
Um yeah, or it may be that that is your only option if you're not in a position to make a very large purchase. Well, this all sounds very exciting. So, is RLC Pro AI ready? Is this uh how can people get it? I'll I'll take that. So, RLC Pro AI uh is in active development. Um but we are already accepting signups for our our tech preview um access. Uh uh so in early June, we're accepting signups now. We're engaging with people. I encourage our listeners to jump in. We want to hear from you and understand your needs. Uh and then in early January, we will um open uh uh up the bits um delivering you an image uh with a major set of the features that we plan to deliver uh for GA and uh then later this quarter uh we expect to be in a position to deliver a GA release.
Um we want to start talking to you now. We want to start getting feedback from you now. Uh we want to start building something that is built to address your dates. Yes. Yes. That's very exciting. Yeah. The tech previews have been going really well. Um it's nice to have people come in and then you kind of get to be, you know, on the ground floor being like, okay, like this is cool, great idea, but like this is how it's working in actual, you know, reality. So, like add this, do this, do that. And it's Yeah, it's awesome to be responsive to our customers and people who are out there doing the good work.
It's a it's a win-win. Yeah. And I'll add, look, no knock on sales. I know sometimes uh uh engineers and actually the implementers of this want to be able to get their hands on technology um without having to engage with a procurement process. So the nice thing about engaging in our tech preview program is you know what you can get the bits, you can deploy them, you can play with them. We're available, our engineering team is available. We can have dialogue and conversations and you can kind of safely figure out um you know run it through its paces and figure out if this meets your needs.
Yeah, that's exciting. Okay, so is there any um key differentiator differentiators? Wow, that's a good word for me right now. Um, that potential users should be aware of from RLC, RLC AI and others. Yeah. Uh, I mean from a right now it is latest. This is you're going to be getting the latest is one of the differentiators. Um this is a an upstream kernel and generally we're keeping it very close to the upstream and but it's still supporting um in a rocky Linux from CIQ9 user space. So it's the same user space you're used to. That part hasn't changed. Um so making sure that your software stack is still compatible in user space is a major part of this and then you don't have to you know go through Linux from scratch to be able to get the most cutting edge features.
Beyond that, we've also um besides the modifications and kernel level optimizations that we're adding around build time and various other compiler flags, we're also adding in additional frameworks, libraries that are used by basically everyone. PyTorch as an example, TensorFlow um on their road map is the Onyx format. Um, and then there's uh hugging face libraries, the MLOps uh components that having them baked in uh in that gold image just drastically speeds up your you know uh time to token because you don't have to sit there waiting for things to install and then you don't have to deal with any complexity around compatibility or dependencies. So all of these things being baked in means uh you don't have to set that up from scratch.
more reproducible, but it also means that you're not going to be getting certain features that are focused on other types of workloads. For example, FIPS 140-3 might not be directly applicable to AI because it's moving so quickly that you're not going to want to wait for, you know, the the code to freeze long enough to be certified. That's great. Thank you. So, I'm I'm really excited. I know that we've already had a lot of interest. people are are asking us about this. They're like, "Yeah, because I want that new driver. I want that that new hardware. Like, I want to I want to run fast. Like, let's go." Uh, so there's been a lot of interest.
It's very exciting. We'll put um the link for to join the tech preview and then of course you can always go to our website and fill out any form and that comes directly to us and we'll make sure that you um get, you know, all the bits and the pieces. And you can see Brian, he's very excited to get this in your hands. So definitely I want I want to talk to people. I want the data. I want I want people to get their hands like you're you you called me out. You were right. I mean you want you want want real people using it because then you know how to make it work for real people.
Yes. Yes. Yes. Yes. That was perfect. I feel like most of you got a little final thoughts. Rob, you got a final thought here for the people? I mean yeah. The core difference between like what I'd say this image is and a lot of what the other images that CIQ produces is we are trying to get closer and faster to what's coming out and being very responsive to the AI market versus some of the other markets we look at without compromising on those things like harding and security. Yeah. Yeah. Let us do the grunt work for you. Yep. Make it easy for you. Awesome. Well, thank you guys.
That was amazing. It was so good to get together with all of you and I have a sneaking feeling that I'll probably be seeing all of you here sometime soon again to continue talking about the cool things that we are doing um you know building on top of Rocky Linux. So first of all thank you to the Roxy Linux community for all the good work that you guys are doing and making this possible and everyone else and all of the upstreams and all of the projects that we are a part of. It is very exciting. people that are working on the kernel so that we can grab it and take it and the hardware and all of the good stuff.
So, appreciate it. Again, go to ciq.com. We'd love to connect with you and get you set up with RLC AI. Well, thank you guys. Have a wonderful day and we'll talk soon. Thanks. Appreciate it. Thank you for being such an awesome host for us. Thank you everybody.
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
9
Enterprise products
Spanning the kernel to the orchestrator
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.
