
Organizations are committing hundreds of millions of dollars to GPU infrastructure and running it on operating systems that were never designed for AI workloads. The OS underneath your GPU fleet determines how much performance the hardware actually delivers, and for most enterprises, that performance has been left on the table.
RLC Pro AI is purpose-built to change that. The CIQ Linux Kernel, GPU drivers, libraries, and frameworks ship built and validated together for AI inference workloads. No manual CUDA assembly. The same validated stack runs on bare metal, AWS, GCP, Azure, and sovereign on-premises infrastructure from first boot.
This session walks through why the OS layer is where GPU ROI is won or lost, how RLC Pro AI is architected to maximize output from the hardware enterprises are already running, and what production readiness actually requires at the OS level, with a live deployment walkthrough and Q&A.
Transcript
Good morning, good afternoon, good evening, wherever you may be hailing from. This is the CIQ webinar series and we are live today as of April 2nd, 2026. Yeah, that's right folks. 2026 is already a third of the way over, but we are just getting started here at CIQ. Uh March was a crazy busy month. We had RLC Pro, RLC plus, some partnership announcements, uh the CLK kernel C3 announcement. Well, technically that was today. Uh but uh one of the big things that we uh we came out with last month uh in the month of March was actually RLC Pro AI. So what we're talking about today is how do you maximize the throughput of your AI infrastructure?
And it wouldn't be a webinar unless I bring in some a few friends of mine. Uh some some folks that are brilliant in their field and uh really enjoy engaging with them. But before I introduce my panel of guests today, there are a lot of them. Uh today I'm excited to share that um we are live. Um and so there is a chat function and uh so if you click on comments uh feel free to say hello and where you're from. Uh if you have questions throughout, make sure that you uh throw your questions in there and we'll do our best to address those as we go along today.
If not, towards the bottom of the uh towards the bottom of the hour. We'll we'll try and cover as many questions as we can live. That said, let me bring in Mr. Brian Dawson. Uh Brian, you uh you finally made it onto the CIQ webinar series with me. I appreciate you joining me today. >> Uh great to join you, Eric. And uh yeah, excited. I end up uh missing a couple. Uh it's it's exciting to uh uh to jump on this new iteration of CIQ content. Uh to take a minute to introduce myself, I am Brian Dawson. Uh I am a director of product management uh within our Linux area within CIQ.
And in particular uh I've been focused on uh our workload optimized variants. And the one I'm excited to talk about today is the one that me and the team have spent months on uh tuning and bringing to market and that's Rocky Linux from CIQ Pro AI often times referred to in short as uh ProAI. Um so looking forward to talking about that with you today, Eric. >> Awesome. Drian, thank you for joining me. Uh also joining us today from CIQ is Mr. Damon Knight. Damon, welcome to the CIQ webinar. Hey, Eric. Uh, wonderful to be here. Uh, everybody, my name is Damon Knight. I am a senior principal automation engineer at CIQ and I'm also kind of our all-around AI nerd.
View full transcriptHide full transcript
>> And nerd among other things, as I found out in a recent podcast interview with you, so uh um but uh I I won't I won't open up that can of worms. We only have an hour. >> And also joining us is our friend Zack. Uh, Zach, you don't work for CIQ and uh, we'd love to find out where you do work and a little bit about yourself. >> Yeah, I'm uh, Zach. I'm from AI Insight Solutions, which is an AI consulting firm. I help startups to medium-sized companies build cool things with AI and machine learning. As Damian said, I'm also a huge AI nerd, and I'm excited to be here today for the conversation.
>> Awesome. Yeah, glad glad you could carve out some time and join us. there's uh you know it's the the technology industry can become a little bit of a of a of an echo chamber. So it's nice to get some some outside opinions and you your your advice has been invaluable as we've been working on pro AAI. Um so uh once again if you are watching us live there is the the chat function so feel free to throw uh throw a greeting out there and then if you have questions as we dive into today's topic u would love to hear your questions. Uh, let's see. All right.
Sorry, I forgot to pull up my notes. Doing great as a host. Um, so first, uh, before I ask my, uh, panel our first question, I'm actually going to ask the audience a question. So, how are you managing your CUDA driver stack today? And so, I'm just going to throw in a multiple choice question there in the chat. Um, so if you are watching, why don't why don't you just read that question and and answer A, B, C or D in chat. Really curious to see how folks are using their their current stack. And while those answers are coming in, uh, Brian, I want to start with you.
I want I really want that highlevel industry focus uh that you bring to the table. And so question number one for today is where are most organizations right now when it comes to GPU infrastructure and getting AI into production? That's not just running, that's in production. >> Yeah. And uh I'll I'll go back a little bit and describe what I've observed after a number of interviews, months of research, um is that what we're finding with uh production is, you know, we've talked about we've been on this technology evolution that is much more compressed or moving much faster than than those who have been at this for a while have really ever seen before.
Um and what that has sort of led into is uh uh people needing to move fast, begin experimentation and move to production very quickly and faster than people have necessarily had the ability to build from ground up what the op uh what the what the optimal AI infrastructure looks like. So what I am seeing uh very frequently is that many people um started experimentation on the cloud right hard to get GPUs easiest place to sort of secure GPUs and hardware to work um they started to build and then quickly move into production and they found okay yes we found some business value um but we've also found um that if we don't intelligently manage spend this is expensive.
Um people have then started to acquire GPUs, realize that they need to move at least some of their uh their their uh their their workloads on prem and then now people are encountering okay I was able to outsource stack concerns to things like AI foundry um I was able to not worry about performance because we were up tuning the model level and now they're in a world where they have to figure out how to how to instantiate or reinstantiate or or mimic um this infrastructure that they started out on uh internally within their own traditional IT operations. >> That's awesome. Zach, is is that kind of what uh what you've seen from the industry as well?
>> Yeah. Yeah. I I've seen a lot of similar things. I've a lot of orgs I work with when the spend question comes up, you know, you look at, all right, how many bucks are we spending per hour? Let's just go buy some GPUs and stick them in a box and put them under somebody's desk. Um, and now you've, you know, you've solved one problem. Perhaps you've got to fix spend, but now you've created a whole another set of problems for yourself. The thing sounds like a jet engine. it's uh heating up the office. Um the monitor stopped working when you updated the CUDA drivers and you know the guy who knows how to fix it is out for the rest of the way.
Um so yeah, it's it's there there are no there are no easy answers. There's a lot of people working very hard to keep to keep their stacks running so they can do their job. Well, and and the one thing that uh that that Brian and Zach missed was then the CIS admin and networking teams are looking at the IP address for that random box and going, "Where is this? It's it's not in a rack anywhere." >> Yeah. Yeah. Well, yeah. The old days of shadow IT, I thought of, you know, my my my my early career, the first iteration of CDW, right, where you just buy a box and ship it under your desk because you didn't want to wait for central IT anymore.
But we're getting a version of that of what Zach said, right? And and that's actually great from a a IML um perspective, right? I need to move workloads over. I'm doing experimentation. Um now this thing is real. I got to mitigate the shadow IT. So now I'm going to our central IT organization, our Linux system administrators, and I'm going, "Hey, I need you to take this over. I need you to rack these cards, rack these machines, metaphorically speaking." And what I'm hearing from uh many of the system administrators is uh Debian, I don't know how to run Debian cards, right? And there's this whole new stack that people have to figure out um sort of how to set up.
I would add uh briefly that, you know, another state that we're seeing is um people have, you know, as as someone's used an analogy for people are figuring out how to drive a Prius, right? I'm just trying to get on the track and do a couple couple of laps and we are now starting to move to the place where um we have to get into performant cars, right? We're now going to take high-speed laps around the track. Um so you see people going, "Okay, look, my Prius was sitting in a container. I got my general structure um going. Now I need to figure out how to squeeze the most out of it, right?
How do I get it to perform? How do I get it to be stable? How do I enable myself to utilize the most of my GPUs that we spent significant money mill money on? >> And I I agree it's, you know, if if if AI is going to prove itself to be a legitimate tool, and I I think it will be. We may not quite be there yet, but I think it will prove to be a legitimate tool. And the the real question though is how much time right now is an AI andor systems administration team actually spending building the infrastructure versus actually using the models that they're building.
>> Yeah, I I'm going to take a lead on that because and then I want to throw it to Damon. I'm just going to say I'm hearing a range. I I will, you know, I'll throw it to Damon who can maybe uh uh uh underline this or or correct it, but um I heard that on a happy path to set up sort of my golden image or one server um that is going to be a couple of hours. I personally have heard and experienced that that can actually um uh span across days um to get your GPU set up and running. Um, so I'm hearing it's been painful for people, but luckily we have our resident AIM ML nerd here uh that has had to do it many times.
>> You know, it's it's not fun. Uh, if I a lot of us, you know, might might remember the the early days of kind of it being a thing, right? to to you know some of your earlier points about things being the wild west and and the AI world still feels a whole lot like that right now where nothing is standardized nothing works the same way there's there's very little abstraction uh and so yeah it it takes a lot of time potentially to even just figure out what you need for the stuff that you've got handy right um you know no one about the time that you're going to spend googling how to install everything when they plan this stuff out.
But but the fact of the matter is actually you're probably going to spend more time googling how to install everything than you are installing it, which will also still be more time than you spend using it at least for for a while. Um, so it's bad. Uh, even even for someone like me who's done this a million times and has written scripts to do it for me, uh, it's still, you know, 10 15 minutes on on a really good day to get a single box up and up and running with CUDA and the right drivers. That's again, that's that's best case scenario where I've already done the work and I'm basically hitting enter.
>> Yeah, I I can I can relate to that. Uh the first time I set up a local LLM, it probably took about 3 days off and on trying to get things to work. And then even after things worked, it wasn't getting the performance that I was expecting. And uh it it was it was pretty brutal. So I I can definitely say that in in my time trying to run local AI, I spent more time uh probably 90% more time focused on getting the thing set up than actually doing anything productive with it. >> Yeah. Yeah. How how about uh how about from the advisory uh perspective, Zach?
How how long do do people usually take to get to that first inference? >> Yeah, I mean uh I think I think Damon put it pretty well. Like everything goes well. You know, it's under an hour. Um you remember what you're doing. You do it all in the correct order. You get all the steps right and there you go. You know, you've got the CUDA drivers, you've got the box running, everything everything works. Um, it never goes right though. Uh, my experience is always something along the lines of, oh, we updated the CUDA drivers. Uh, this literally happened to me once and now the monitor doesn't work.
Um, so you can still SSH into the box. Um, but good luck if you want to, you know, connect a keyboard and mouse and a monitor to it and actually like log into it, you know, the way a normal person would with a computer at their desk. Um, so it's it's it's all over the place. It's always uh it's always an adventure you're going on whether you want to or not. >> For sure. >> I'd say and you think about um you know um when we're all excited. I think about a very personal experience, right? Uh whether you're talking rack scale or you're talking my Dell.
I'm excited. I got my Dell. I'm ready. I'm going to load some models. I'm gonna I'm gonna run. Next thing I'm I'm I'm down in grub configs. Um blacklisting driver. >> Right. I'm I'm and I'm taking photos with my phone with chat GPT so I can say what what is this screen and how do I get to the place where I need to change the thing I'm trying to do. Yeah. >> I think you guys have been watching me as I work in my home lab. >> So here's here's the problem. And an hour or four hours may not sound like a lot but the problem is I don't have just one AI box.
I have like 50 of them all with multi,000 cards. As as a systems administrator or AI engineer, I don't have that kind of time to spend on each box fighting things through. What I mean, why not just uh why not just take an AI workload and put it on a typical server? I mean, that that works out fine, doesn't it? What do you think, Damon? I mean, you know, if if you don't uh mind the job outlasting the the heat death of the universe, then then that's probably a good plan. Uh well, you know, it's AI workloads. Uh and I know I'm preaching to the choir a little bit here, but you know, AI workloads are very different from what historically data centers and and data processing folks have really dealt with, right?
Um, unless you were doing something special with, you know, video processing or whatever where you needed, you know, basics or or some sort of special hardware, things generally have all just been running on general purpose CPU, you know, utilizing that that processor, utilizing your system memory, all of that. And so your your OS spends its time managing the CPU and the memory and those resources. And when you start going into AI workloads, um really the CPU and the system memory fully, right? You you've got that big fancy GPU that you've, you know, spent more than a house on and you you want all of the processing to stay there the whole time.
You don't want to be offloading stuff. You don't want to be swapping stuff around. You don't want to be communicating or spending kernel time. Uh and so that's a a really big shift. Um also, you know, scaling becomes a whole lot harder, right? There's there are hardware constraints. We we all know the difficulty of getting GPUs. Uh you know, just just have more, just throw more at it. These these options all kind of go out the window in in this brave new world that we find ourselves in today. >> Yeah. I I I' Yeah. Risk to add.
Um it's interesting that you say right like when you're talking general purpose I need a server that's a mail server I need to run a database right um the whole idea of vertical versus horizontal scaling um is just different in nature right um than how do I horizontally how do I horizontally or vertically scale an AI workload right um um uh I'm generally looking to take a general purpose Linux OS um deploy it to a serve as many use cases as possible and occas occasionally fizzle fiddle a couple of bits. My job is based on this server is stable. All right. Um so yes, I have all the services running.
I'm making sure that my IRQ or my interrupt prioritization is right. So everybody gets attention. As Damon says, it's a very different um um um um universe for AI. And look, the reality is is um um for understandable reasons, uh people weren't too concerned with that. They were driving their Prius around the track. Right now they have to figure out how to get it to act like a Ferrari. Um um but the reality is something like 52% of AI workloads are actually running on a general purpose OS. >> Yeah. So, we we kind of hinted at a couple of the different steps, but Damon, can you walk us through if I do take my general purpose ISO, I just go out to, you know, rocky Linux.org and I grab the community edition.
Uh what what's kind of that process to getting from from nothing to something? >> You know, it's uh not easy. Uh generally speaking, right, you're you're going to be making sure you have your Linux kernel headers and and development libraries installed. uh you have to go find the right just straight up normal Nvidia driver versions um that support your hardware um which is not necessarily as easy as you might think. uh you know you have to go find the right CUDA version. Um, you might be going, Damon, you know, the drivers ship with CUDA. There's CUDA and then there's the CUDA toolkit. There's a whole separate thing that you need uh if you're going to be dealing with AI and ML workloads.
Um, I' I've told this story to to Eric before, but I'll tell it again. You know, if you uh if you're an Auntu user and up until just recently, uh, you know, 24 was the LTS and 25 was kind of the the bleeding edge. Um, and if you ran abuntu 25 and you went to the official sources, there was no CUDA toolkit for Abuntu25. You had to know that you were supposed to install the the Iuntu 24 CUDA toolkit. Uh, good luck to you if you didn't. Uh, but you needed to know that, right? Uh, and you know, then, okay, do you have the right torch installed?
And was it compiled against your CPU or the GPU? I sure hope it was the GPU. uh otherwise you're not actually using it and and you have a big expensive uh power- hungry paper weight. Uh you know there's there's a lot of steps involved to get to actually you know taking that OS and and getting to tokens being >> generated. So, Zach, is is that kind of what you found with some of your clients or any kind of hidden costs that pay? >> It's like that meme of the guy with the whiteboard and all the strings like you're you're in a room somewhere trying to write out the matrix of, okay, we got this card and it'd be really nice with monitor work.
So, which driver will do both, you know, both of those things with the CUDA version we want and also, you know, I'd like my monitor to work. And then the torch drivers, the two CUDA cool I mean I I can't tell you how many conversations I have with IT departments where it's like okay great uh the CUDA drivers are on the you know the the the drivers are on the machine you should be good to go and then I come back I'm like yeah actually I need I need CUDA too. Okay couple days go by right we got CUDA too. All right great uh I actually need the the the CUDNN libraries as well.
Do you have those? Oh, wait. No, they're not compatible. Like, and so you're back and by the time you build the stack up to right now I'm at torch with the right version of the toolkit with the right version of CUDA with the right version of the drivers with the right GPU. Uh, you know, how many how many back and forth in that cycle have have gone through? How much just human time have you burned like getting all the version numbers lined up and everything talking to each other and plumbing your data from, you know, torch at the Python level down to your GPU.
Well, and not just the human time, but the trade-off time is brutal because you know that your your infrastructure team goes in, they fix they they fill a request, they close it, they move on to the next thing, and then you come back a couple of days later and go, "Oh, I I need this this extra thing or this extra thing, and then it takes a couple of days to get through that process again." So, the the handoff time is just murder when when you have to keep going back to it, >> right? Yeah. I mean, and it's just and you're speaking different languages. Like you've got, you know, data scientists and ML people talking about, you know, CNN's over here that they want to run because they're doing an image problem.
Uh, and you've got someone in the IT department over here who's like, you know, I don't know, looking at a thermal graph from a server and saying, "This thing's going to melt itself." >> So, I I feel like there's almost there there's got to be a better way to do this, Brian. I just I feel like if if only some company out there somewhere had a cool idea and and fixed this problem. >> You know what I think would be really cool and this is just you know it's great idea here. I think I think what would be really cool is if someone was able to simplify all of that dependency management, validation, kernel mod building, tuning of PyTorch flags, and then packaging it.
That would be awesome. So that I can just get an ISO or a cloud image and I could hit install and then I can download O Lama and I can get some tokens out of it. That would be really cool. Um and in fact that is what um um this team here at C CIQ has done uh in terms of uh building delivering uh uh Rocky Linux from CIQ ProAI um is uh build a uh image uh or a variant of our trusted Rocky Linux that will enable people from Damon to Zach and even me uh to be able to grab an image, put it on on a box and be able to be getting tokens returned from an inference server in as little as three and a half minutes.
Um, uh, uh, so I think we've solved that problem, Eric. I think we've solved that problem. >> So, Damon, I I want you to kind of tee up, uh, the demonstration that we have going because this is going to take some time to run and I don't want to spend a good chunk of our hour together. Uh so if you can can kind of tee up the the longer version of what what we're getting ready to to show here. >> Yeah. Uh so you know I've talked a little bit about uh you know kind of that happy path that best case scenario where someone has a script ready to go to do all of the things.
Uh and and I do have one of those scripts and we've got a couple of systems. Uh we have a uh Iuntu I think it was a 24 system. uh maybe 25, I forget. Uh and we have an RLC Pro AI system. Uh both with identical hardware, identical configurations. Uh A30 GPUs, I believe, is what we're going for for this demo here. Uh and just kind of really helps you all see the difference in completely fresh install, right? Same state for both to that kind of time to first token, first inference. Uh, and I think we're we have a little demo that shows uh it talking about um AI in fact.
>> Yeah. And so I'm I'm just going to I'm just going to leave this here. We're going to let this run. This is the the Ubuntu system uh that that Damon was talking about. Not not to not to pick on Abuntu, but it's just uh it's what we've been testing with. Um so well well that's >> if I if I can too for for folks you know we picked Abuntu because honestly right in the AI space historically that's where folks have gone first right because of the the kind of community support and and all of that. Um so we really wanted to make sure that we were comparing this to the thing that most people are using today.
So Brian, I I wanted to circle back now now that we've now that we've kind of got our cooking show uh set up here. We're going to let we're going to let Abuntu cook for a little bit. And so while that's cooking, I I want to circle back to something that you said. You you you use the term validated and I think that's really important and I want to call that out. There there's one thing to just go in download a bunch of drivers on a version that you think will work, but validated is a whole another level. you wanted you wanted to talk about uh validation a little bit.
>> Yeah. Uh and if I may, it's actually a good time for me to pull back up a bit and and just kind of give a once over of uh RLC Pro AI and what it is. Right? We've talked about some of the problems that people are experiencing setup time that uh that that has opportunity cost. It has, you know, up to $3 a GPU hour hardware um sitting there. it uh we have performance left on the table right as we all just talked about. Um so what uh RLC Pro AI is is it is a AI workload tuned uh uh version of Rocky Linux and effectively Rocky Linux from CIQ uh that enables people to set up a system and get it getting it running in as little as minutes.
It pulls forward the latest upstream stable kernel where you will find that uh vendors are committing uh support for uh current or emerging hardware so that you can get the most out of those hard that hardware's features uh and you know it's uh it's enterprise Linux uh so it is enterprise Linux uh such that people that have invested their time working on in real and have a need for stability in the enterprise are able to get the balance of speed and stability. Now we get to where we lean in to validate it. One of the things that we have done and as Damon sort of called out uh I think we learned during the process is more painful than I ever would have assumed right is we've pulled in everything you need in the stack right and we started with an initial stack and we will be expanding that.
But right now that means we pulled in the latest upstream kernel. We brought in all of the K mods for Nvidia and then soon AMD as we've announced our partnership. We've brought in the CUDA toolkit. We've brought in some of the other user space requirements. We've brought in PyTorch. We re we've recompiled PieTorch, benchmarked it, and tuned the flags. I you know, and switch flags to make sure we were getting um uh uh the most out of the GPU throughput for production inference that we that we could. Um now, that's great. We've put all of that together. But the keyword validated, okay, throw in PyTorch, throw in a KM mod, throw in the latest kernel.
Um, next thing you have to do is make sure that that all works together. And to something Damon indicated earlier that, hey, great, we didn't throw this stack together and just kind of get it to work, but we've left all your performance and utilization on the table and your systems running hot. So, what we've done is we've validated that from a dependency standpoint, these things work together. We've also validated that from an operation and performance standpoint that it is performant. >> So Damon, what what can happen to a system that hasn't gone through this validation process where it is just kind of duct tape and hopes and dreams that are are keeping all the version numbers aligned.
>> Well, you know, I I did mention a little bit of it earlier, right? Uh but you know a really I think somewhat frighteningly common one is you know that that whole uh torch piece right it's really really common to wind up not running your GPU enabled torch like you think you are uh and to actually be sending all of your AI workloads to your CPU >> and you know you will pull up you know Nvidia SMI you go why is my GPU sitting at 0% utilization and why am I generating half a token a second. This this seems odd. Um but that's you know a really really common occurrence.
Um you know Brian mentioned that you know we compiled you know torch um with flags right? Uh when you're running inference you're going to see all kinds of messages pop up for oh hey you know you don't have flash attention compiled so we're doing this in a you know a silly slow way. Again our our validation our testing we we enabled these flags. We've run the benchmarks. We've run the tests.
We've seen that this stuff actually is working together uh and is using the right optimizations and setting all the things to the right places so that you don't waste your time again looking and going, "Boy, boy, where are all of my tokens going?" Well, and I I love the fact that that the engineering team has spent so much time testing out these different versions, introducing these different flags, but it it actually goes a lot deeper than that in that we aren't shipping the standard Rocky Linux kernel. U so part of our announcements in March was the CIQ Linux kernel or CLK kernel. Brian, you want to talk to us about about what the CLK kernel is?
>> Yeah. So what the CLK kernel is uh is it is uh the it's we'll call it out of tree uh right for and I think many of our audience probably knows this um but traditionally if you're uh uh enterprise Linux distribution uh that is downstream of Red Hat Enterprise Linux or RE uh through that overall stream um you are pulling your kernel from kernel.org Now enterprise Linux traditionally and rightfully is very focused on stability right we are running business critical enterprise workloads so um we intentionally in that space don't um don't rev too fast so we are careful about um the timeline it's different with with each distro right at which um you pull forward uh uh kernels from upstream kernel.org.
So what is click? Click is CIQ Linux kernel. Currently it is at 6.12. Uh we will shortly in a couple of weeks here be releasing 6.18. Um and what it is and and what that is is that is as fast as we can follow the latest stable upstream kernel from kernel.org that we uh bring down uh we rebuild and we sign. And then we vendor back and support it, right? Validate it, bug test it. Um, and it is vendorbacked. Um, and again, as I said earlier, what is the value of that is because now you're in a situation where you're getting the latest performance enhancements primarily in terms of hardware support and for where many of us are concerned, our GPUs, but also in terms of IO scheduling, um, storage scheduling, right?
It's under continuous improvement. So what we're able to do with with with sort of our deep um enterprise Linux experience um as well as pulling forward and productizing this thing we call CLK or click is um we're giving a balance of the stability of enterprise Linux but speed that gives you a performance advantage and even more so enables you to increase your GPU utilization by leveraging advanced features. um day one out of the box by as much as uh 30 to 50% going off an industry stat that says uh most GPUs are underutilized by as much as 30 to 50%.
Now to be perfectly honest that's for various reasons right it's your data piping it's your pipeline it's your kernel um but but what we are looking to do is take a material chunk out of um uh properly supporting your expensive hardware so you are getting more utilization >> awesome so Damon uh why don't we jump back towards our demo here and kind of walk us through uh what what's been going on over the last eight or so minutes that uh the the Abuntu box has been plugging along. >> Yeah. Uh we've we've made quite a bit of progress actually. Uh we've gotten the Nvidia drivers have been installed.
Uh CUDA has been installed. Uh at the moment I believe it is doing all of the torch uh setup and installation. >> I see my CU DNN libraries installing. That's great. >> Yep. Yep. Uh so I think uh we're we're getting pretty close to the end here. Uh next up I think is going to be doing uh installation of transformers and some other things. Uh we'll see what I what I type. Yeah, look at that. Transformers. >> Uh then once these install, I think uh it's finally ready to uh kick off the actual demo and start generating some tokens.
Is this a case like you brought up CDU and um is this a case where uh and I hear this so I want to ask you as a a IML engineer you ask for your build or your image you get it and uh the framework or library needs not there and that's that trip that Eric talked about where you're going >> right right in the middle there and so someone's gone through you know like this is everything worked you know but it it probably didn't take eight minutes to put this stack of commands hands together and figure out what do I run in what order
and so someone's got to go back and rerun this process and insert some steps in the middle and that's not always as easy as oh just just running it again you know um so so yeah it's it is there are a lot of parts to put together it's very it's easy to forget one um it's like the IKEA couch you know you put it together um and there's a couple of screws left on the table and you hope those those weren't the important ones that Uh you can see here actually we've we've finally uh we have moved on to the actual uh generation demo. It's just uh loading all the models up now.
So we're any minute now we're going to start seeing uh actual tokens from Abuntu. >> So how how long do you think it took you to build uh to build this demo to to Zach's point, Damon? Oh, I I don't know that I could even begin to guess. Uh, this this lovely script uh was a product of months and months and months of uh me doing more installs than I care to remember. I'm pretty I wake up in the middle of the night screaming cuda in a in a cold sweat. >> Yeah. So just just copy paste alone, we're we're up around 12 minutes here and we're we're still we still haven't gotten to ask any questions of our of our AI model.
So we'll we'll let that run a little bit longer. Um so why why don't we uh why don't we go ahead and talk about uh what what comes next? Uh so it's it's fetching files. It's it's kind of uh it's doing some downloads. What what can we expect next, Damon? Well, I said next uh we're going to see actual generation. Um the script is is written up here. Um so basically it's it's running a very very small uh Quen model. Uh and we're going to ask it about uh LLMs I believe is is the question that we ask it about. Uh and we're going to see you know some some answers and some information being generated for us here.
Um, but you know, uh, the joy of AI, I'm sure you all know, uh, is downloading those models. You can see here, this is a 16.4, uh, gigabyte file here. So, it does unfortunately take a little bit even even on a pretty fast connection >> and it's that's a relatively small one. I I joke with my friends and people ask me what I do for work, it's well, I watch I watch uh progress bars on the command line in many different forms. I've unfortunately I've I've been playing with some of the the very large models that are you know 400 500 gigabytes in size and yeah you start the downloads and it's go for a walk but okay here we go.
>> It is generating some tokens >> and I will have some questions Damon about kind of what preceded this when when when we get a moment. So, so Eric, uh, what's our what's our timer at here for for tokens? About 13 and some change. >> So, we, uh, yeah, so it's, uh, 13 minutes and, yeah, just a bit of change the seconds scrolled off the screen, but uh, yeah, about 13 and a half minutes. And that's having a a copy and paste procedure. That doesn't count the months that Damon spent beating his head against a keyboard hoping that the right magical incantation uh appeared on the bash uh prompt.
>> Well, can I and can I ask because it's important here in the time that we spent on it. We say copy and paste like in the in the real world someone may be copy and pasting from a tutorial, but tell me if I'm wrong, Damon. Uh more specifically, what the video we're watching was a script that was run. Is that right? Okay. uh effectively uh I was actually copying and pasting. >> Oh, you were? Okay. That was my mission >> up until that final uh you know demo uh which is a script itself. Um that was me copy pasting commands uh and typing a few by hand uh when I was feeling spicy.
>> Okay. When you're feeling spicy and I'll say just you know as we talk openly because I look and I go wow that was even faster. Like someone may sit here and go yeah you know 16 17 minutes not horrible. Um, but as you call out, Eric, there's a lot of sort of learning that was into it. But Damon, I also wanted to chat like when we started, we immediately went into a reboot, right? So, there was at least one boot cycle in that. >> I think there was three. >> There was three. >> There's there's depending on the exact system you go at. I don't I don't want to mislead anybody.
Uh but there can be anywhere from zero to three reboots depending on if you have to and and you know to to jump right into it right that very first thing that we did was blacklisting Nuvo and rebooting right if you've done GPU stuff you know that uh if you're dealing with local hardware you are basically always going to have that step uh you know if you're if you're in the cloud if you're in you know Azure AWS you might get spared needing to to do that neuvo blacklisting, but if you're on prem, yeah, there's there's a reboot that's going to be part of it. You have to, you know, uh regenerate um you know, an RAM or DRI potentially if you're dealing with secure boot and you're going to have to potentially install new drivers that have to be signed.
That will be a whole other reboot right there. Um >> so it can it really adds up. Yeah, because I'm sitting here going, man. I mean, like, compared to what I experienced, that's not bad. Like, am I just really that inadequate? Did I miss something? I don't expect you to testify that I'm not, but I wanted to >> for for what I wanted. I I wanted to present everyone in the best light, especially Abuntu. So, right, that was a nonsecure boot system so that we could just install drivers uh and all that without an extra reboot.
um you know RLC ProAI if it had even if it was a secure boot system it would have worked out of the box no reboot needed right with with iuntu no secure boot we get to skip one of the reboots but if there was secure boot that's that's another reboot that you know uh an IT team or cisenge has to to worry about and deal with and coordinate >> okay if I may one more thing Eric sorry you there DKMS was in there you you compiled Kods as part of that process uh for the the Abuntu stuff uh it they just kind of abstract and handle all of that.
Uh again, right, unless you're unless you're dealing with the secure boot ones where you need the the signed packages, there's there's not a whole lot of rebuilding uh that you have to do thankfully. So that that was pre-recorded uh to be transparent and that was months of work and I I got to go on I got to sit inside car as Damon went on this journey uh and there were times where there was m much yelling at clouds and and uh and long long nights but uh I I wrote it down at 10 minutes and 32 seconds is when the demo switches from getting all the prerequisites setup to actually starting to work on our AI models.
So, with that in with that in mind, a 13minute demonstration, 10 and a half minutes was just setup. We've got another system here that uh is focused on RLC ProAI. And our first step was not to blacklist Nuvo, but it was to actually just check and make sure that oh look, we see our card Nvidia SMI is working the way it should, which in my experience is always a challenge. U but we just kind of dive right into now we're doing pip installs, right, Damon? >> Yeah. Uh it's great. We not only are we going straight to pip installs, but we're going straight to transformers and accelerate, right?
We're not dealing with that long torch build and and compile. That is what takes a really good chunk of time. We get to skip really far ahead. We're a couple of seconds in and I'm already running the the generate demo. Uh there's no getting out of downloading those models. I'm sorry guys. If if I could preload all the models in the world, I would. But you're you're >> But then there'd be a new model tomorrow and you wouldn't have it. >> Exactly. Exactly. >> And your your base image would be terabytes in size. >> Yes. Uh but this is this is the exact same demo, exact same code um just running on our uh ProAI box.
>> So how much Zach, how much does this change things for for your your clients, for yourself? Because I know you're you're an avid AI user yourself. >> I mean, um my time is very valuable. You know, it it is funny like watching that first demo. Uh this is actually the first time I saw it. Um, and I see a lot of things that are familiar. And it's funny because it actually felt very good watching that because like this is this is everything going right. Um, you know, this is this is the day you wake up and it's like you show up at the train station right as the doors open.
like everything just takes perfectly and it's like I I' I've had that day occasionally where where you start from a cold boot and you just go through the script and everything just works and you end, you know, and and 13 minutes later you're running your model and that's that's awesome. That's great. Um it's really nice to not have to go on that journey every every single time I've got I've got a new machine that I want to start. I I will say the the story that I'm sure Eric will enjoy is the story of filming that demo. Gosh, >> you know that that involved me rebuilding the whole box from scratch like three times as I ran through it and did it wrong and right.
>> Okay, nuke the whole box, start over, re-record. Uh until I got that script into that again, that that nice punch list of okay, cool. I can just do this, do this, do this. Um, it is it is a process. >> See, your your demos required hardware interfaces where a lot of mine are virtual machines. And so after every few steps, I take a virtual machine snapshot just so I don't have to do that. >> So, so here we go. We've already got our our tokens already being generated. >> We we are sitting at 3 minutes and 2 seconds. And and Damon, can I ask because this is also the first time I've seen versions of this demo, you know, so I'm not acting for our audience being that Damon and I work together.
Um, but I am curious. This appeared to actually be outputting tokens faster. Is it possible that that's the case here or is is that an illusion? >> Uh, I'm I'm going to say that's probably just an illusion. They're they're probably roughly the same. Uh but you know with everything else moving so much quicker it >> even that last little bit right it it >> the well but like I think that's important to call out too right the the user experience just feels nicer it feels faster um because you aren't spending so much time sitting around waiting for all those prerequisites to to build compile and install.
>> Yeah. And I mean the model download time is an especially painful one. like you can't avoid it. But it's very painful to realize you've got that model downloaded and you're like, "Oh, I got to wipe this machine and start over." And then you're like, "Well, I don't want to wipe this machine. I want to keep these weights around." And you kind of get yourself into this, well, you're not quite resetting the state of your machine to reinstall and you can get really really trapped. Like it's it's nice to have your iteration loop be you download the new model. You do inference, you look at the results and you can iterate on what you care about, which is which is the AI, which is the models.
>> So Brian, I want to direct this next thought to to you now. Now that you've kind of seen the the finished product, the finished demo that that Damon's been working on, how do you feel that this this scales if if you go from one node that we are working with to to 50, how how does this change things operationally? >> Yeah, this that is very interesting. Let me start with let me ask one thing for me since I am I am I am admittedly new to this iteration of the demo. I know you guys had your stopwatch out. What was what was the time on that?
The the last the on the ProAI that we just saw >> the ProAI video was about three and a half minutes. >> Three and a half minutes from uh fresh install to inference. >> Yep. >> Okay. Um um yeah. So, uh and feel free to, you know, correct me, jump in or resteer me. I'm going to say one thing about scale that is important in what we designed for here.
um in terms of choosing the open- source drivers uh versions of things that we pulled flags that we configured in the CUDA stack is uh one of the things that this does um for you at scale is we've designed this to support the widest range of hardware as possible and you asked earlier about infrastructure right one of the things that is happening that we're hearing about is um is that you know people have to beg borrow and steal whatever GPU they can get and there's often times that if you go back to an amper generation GPU um it can work just fine um for certain um workloads.
So, you know, how do I at scale subvert support a diverse fleet, right? How do I support my Blackwell? How do I support my Hopper? How do I support my Ampere? Um, and a lot of cases I'm hearing about people having to go through this process for each one. So, one of the thing that enables you to do at scale is have one golden image for all of your GPUs, right?
with mitigation of configuration issues you solve on the configuration uh on the configuration um time and then that'll go to um the next thing right um um when I have a validated stack when I have a validated stack that supports the range of hardware my older to newer and I get to use it um then when I do scale that out whether it be via IA pixie boot or what have you when I do scale that out um um I am saved from the inevitable configuration errors that I'm going to find as I go across unique differences across um sort of each of those nodes in my fleet and each of my service.
And so I'll just summarize and I think there's other points or maybe other aspects to that question and to the answer but I'll say it it enables you to scale um confidently and safely mitigating the time that you have to spend iterating with your aim engineers your your your your S sur team um and debugging configuration issues. >> Zack any thoughts to add to that? Yeah, I mean it it's nice to be able to use your old GPUs. Like I you know, you've got you've got a machine sitting around with some uh some consumer cards in it. I I forget even the series like the the 20 what is it 2048?
I don't even remember the numbers, but there's tons of people who who have them. They were cheap at one point. Uh you can put four of them in a box and do some interesting and decent stuff with it. Um, and there's a lot of use cases where you're running smaller models, maybe not even language models. You're doing, you know, interesting stuff with Torch. A lot of the image models are a lot smaller. And it's nice to be able to use your hardware. Um, and sometimes you get to the point where it's like, well, the drivers are old, the machine's old, we can't figure out how to update things.
So, you know, we're just not going to use those cards anymore, and everyone's going to have to just make do with what what we can get for for newer hardware. Um, so it's it's it's nice to let the useful life of these things extend out because yeah, you know, one year it's it's the absolute cutting edge. It's what you're running everything on, but as time goes on, you know, you have you have these old hardware pieces that accumulate. You want to still be able to use them. They're they're still totally functional. Um, you know, um, but maybe the new software stack doesn't support it and so they end up in in, you know, in the trash and like that's just wasteful.
Um, not to mention inefficient. >> David, any thoughts? >> Oh, you know, um, many many thoughts. Uh, and hardware support's always a challenge. Um, Nvidia, uh, and their infinite wisdom, uh, like to change how things get supported and their drivers for different cards, right? Uh, you know, we we could tell so many stories about the open drivers versus the proprietary drivers and the data center drivers versus the consumer drivers and grid drivers and there's all of these different configurations. And, you know, uh, heck, if you know, if you want to run really really old cars, right, you want to run your, you know, your old Voltas and and Teslas.
I think they finally dropped Volta support. Uh, but the Teslas and things like that, right, you need to know which things you you got to grab. Uh, and and it's always changing, right? Nvidia is always updating that. Um, and so it's nice to have um someone else worrying about that. I think uh you know it's even if it's not uh going to cover every single possible use case, um every little bit helps uh every bit that we can can extend the the life of that hardware and the the usefulness of it. Um and again make that that cisenge life a little bit easier. uh you know half of the the folks have very um heterogeneous fleets with all kinds of different hardware.
Uh and so you know whatever we can do to to make that a smooth easy experience is I think just a win for everybody. >> So if I wanted to try RLC Pro AI out today, where would I go to find it? >> So that is you read my mind. Thank you. Um um um so you know a transition from scale is uh you know look if you're deploying uh production inference at scale um I will contend all the stuff we talked about is going to enable you to save time configuring and deploying it is going to enable you to get more out of your GPU and use more of your fleet right um and we'd love to talk to you if that's the case right if you're deploying 10,000 you know 10 thou if you're lucky enough right?
Uh to have 10,000 GPUs uh across x number of nodes. Um but the reality is a lot of us are still on on a learning journey. So I'd also say what we have built in the easy setup time um is just as good for the researcher uh the individual user that is at the start of their journey. And that leads us to where we get this right. Obviously um uh we are in the enterprise business. Um you can always just reach out to us. But what is awesome is one of the other things that um the team has built and the team has announced this year is uh um developer licenses uh as well as trial licenses.
Um so you can go to portal.ciq.com. There you'll see a range of our fantastic offerings that we have and there you will be able to see pro AI um and get leads to securing it of course uh in the cloud across all major 3CSPs uh as well as getting uh qcals ISOs um and soon containers and for all of you who are interested in exploring and uh validating what we've talked about here today um you could simply um sign up for a development license download that and use that. So portal.ci ciq.com. >> Awesome. Yeah, that is the onestop shop for everything that you need.
There's links to documentation, links to contacting our sales teams, links to your tickets, and of course, all those juicy, delicious downloads for not just uh not just Rocky Linux, uh not not sorry, not just RLC, but uh a lot of our other tools as well. Um so that brings us to the top of the hour. Uh I really appreciate the three of you joining me today. This wouldn't be as much fun if it were just me talking into the void. Uh, so Brian, I really appreciate you joining me. Uh, Damon, actually you and I will meet back here in a couple of weeks because just announced a couple of hours ago.
I was so excited to get this email because now I can announce it live on air. But in two weeks from today, we'll actually be back right here, uh, Damon and myself, and we'll be talking with Microsoft. Uh, so we've got a gentleman from Azure. His name is Hugo. He'll be joining us. Uh, wait, Brian, are you on that one? >> I am not. I am not. >> Okay. I might drag you in anyway. But but I know Damon and I will be back with with Hugo and we'll we'll be talking about RLC Pro AI and what that looks like from a cloud perspective. We talked a lot about bare metal today.
Uh, but let's face it, GPUs are hard to come by. And sometimes from a hardware perspective, it's just easier to let someone else do it. And I can think of no one better than Microsoft Azure to deploy uh to deploy those those solutions. And nothing better than RLC Pro AI to run my AI models. I will by the end of the year be using AI to codem a Dungeons and Dragons campaign. That is the goal and I'm using all of our technology to do it. So, it's going to be really great. Really excited. Uh, Zach, thank you for joining us and we'll we'll have to bring you back on uh real soon.
Really appreciate your your input and your insight, especially as someone who doesn't wear CIQ on their badge. >> That was very happy to be here. I love the conversation. >> Yeah, it's a lot of fun. Um, I I I should have teased Brian more, but you know what? What can you do u, but for those of you that joined us live, really really appreciate it. Appreciate the conversation in chat. If you didn't catch us live and you're watching this on YouTube, there is a comment section down below and I check those a couple of times a day. So have if you have any questions, if if you have any comments or insights, feel free to put those into the questions below or into the comment section below.
If you are catching this after the fact, make sure you hit the like button. That kind of helps juice the uh the algorithm. And subscribe because we're we've got a pile literally of about 15 videos that are almost ready to publish. They're they're in final uh final editing. So there's a ton of demos, thought leadership, product overviews, all kinds of cool stuff coming to our YouTube channel. So definitely uh subscribe for that. And uh yeah, so got new videos, new webinar. Uh and then a few of us will actually be at Linuxfest Northwest in a couple of weeks alongside our friends over at Rocky Linux as well.
So lots to do, lots going on. >> Also important, catch us at Human X next week. May be there. We'll be out on the floor. Um uh and we'll be happy to show this to you. >> Yeah. Yeah. Uh I think Brian and Hope uh one of my counterparts will actually be at HumanX next week. So uh if if you'll be there, make sure to stop by the booth. We usually have some fun swag and of course a team of experts to answer all of your questions. With that said, on behalf of my guests, Brian, Damon, and Zach, I've been Eric the IT guy Hendrickx. On behalf of CIQ and our webinar team, really appreciate you all tuning in, and we'll talk to you again real soon.
>> Thank you. And thank you guys for helping build this.
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.