First Boot - First Inference | The IT Guy Show 020
In this sponsored episode of The IT Guy Show, host Eric Hendricks talks with Damon Knight, senior principal automation engineer and in-house AI specialist at CIQ, about why standing up Linux for GPU workloads costs far more time than anyone budgets. Knight describes an industry where hardware is scarce, organizations run mixed fleets of older and newer GPUs across multiple operating systems, and admins spend more time on infrastructure than on the AI applications they set out to build.
The middle of the conversation digs into why general-purpose Linux struggles with AI: workloads are GPU and memory bound rather than CPU bound, CUDA has no simple install path, versions must match the OS and drivers, and Torch compiled without the right CUDA support silently falls back to the CPU. Knight shares his own experience with Ubuntu 25 requiring the 24 CUDA packages and a 13-minute best-case scripted install that many teams never achieve.
Hendricks then walks through how RLC Pro AI fits the enterprise Linux lineage, from the CIQ Linux kernel (CLK) and tuned userspace to drivers, CUDA, PyTorch and torchvision that ship pre-installed and validated to work together. The episode closes with RLC+ NVIDIA for home labbers, thoughts on unified memory systems for local LLMs, and Knight's advice to research ahead and not be afraid to break things.
Key takeaways
- GPU scarcity forces organizations into mixed fleets of old and new cards spread across clouds and operating systems, each with different driver stacks.
- CUDA versions depend on the OS, drivers and hardware; Ubuntu 25 users must install the 24 CUDA packages because none exist for 25.
- Torch compiled without matching CUDA support quietly runs inference on the CPU, showing up as tokens per second far below expectations.
- Knight's fully scripted best case for manual GPU setup is 13 minutes; without a script, 30 minutes to several hours is normal.
- RLC Pro AI uses the CIQ Linux kernel and ships drivers, CUDA, PyTorch and torchvision pre-installed and tested to talk to each other.
- RLC+ NVIDIA gives home lab and community users a Rocky Linux base with the NVIDIA stack already installed.
Questions this video answers
What is the difference between a pre-installed AI stack and a validated one?
Pre-installed means the drivers and CUDA are present and should work. Validated means CIQ precompiles components with the right flags and tests them together on real hardware such as A30, H100 and RTX systems, so Torch is not accidentally shipped in CPU-only mode and every piece talks to the others properly.
Why is general-purpose Linux a poor fit for AI workloads?
AI workloads are GPU and memory bound, so bus speeds and accelerator devices matter in ways general servers never cared about. NVIDIA's CUDA requires special compilers and libraries that must match the OS and hardware, drivers get deprecated, and general-purpose installs were built for uniform servers, not purpose-tuned GPU nodes.
How does RLC Pro AI relate to Rocky Linux and RLC Pro?
Community Rocky Linux rebuilds from RHEL binaries; CIQ then builds RLC+ and RLC Pro from it. RLC Pro AI is a purpose-built variant on RLC Pro with the CIQ Linux kernel, changes to userspace and a pre-installed, validated GPU stack. It remains binary compatible but not identical to community Rocky Linux.
About this video
In this episode of The IT Guy Show, Eric sits down with Damen Knight, CIQ Sr. Principal Automation Engineer, to talk about something every AI team knows but nobody budgets for: the hidden cost of configuring Linux for GPU workloads. Together they dig into why general-purpose Linux distributions were never really built for AI, what actually happens when you spend 30 to 60 minutes per node doing manual CUDA setup, and why a "pre-installed" stack and a "validated" stack are two very different things. From driver conflicts and framework dependency failures to the CIQ Linux Kernel (CLK) and NVIDIA authorization, this episode covers the infrastructure decisions that determine whether your AI program ships fast or gets stuck in configuration hell. Whether you're an ML engineer tired of fighting CUDA every time a kernel update drops, a sysadmin managing a growing GPU fleet, or a tech leader wondering why production AI is harder than it should be, this one is for you.
Welcome to The IT Guy Show, your go-to destination for all things tech with Eric, your friendly neighborhood IT guy! Join Eric as he shares his wealth of knowledge, insights, and experiences gained from years of working in the IT industry. From troubleshooting common tech issues to exploring the latest trends and innovations, each episode is packed with practical tips, skilled advice, and engaging discussions.
This video is part of the RLC Pro AI playlist. Browse every CIQ video by product and topic.
Transcript
Hey there and welcome to the IT Guy Show. I'm your host Eric the IT Guy Hendrickson and this is episode 20. Today we're going to be talking about how can you avoid the chat GPTs and the Claude's of the world and build your own locally hosted LLM. Um we we've this this is kind of exciting because this is kind of our first officially sponsored episode here on the channel. So I'm excited to to share a little bit about that later on. But it wouldn't be an episode of the IT Guy Show unless I had an amazing guest here to to tell me all the things that I don't know.
And so let me bring in Damon Knight. He is a co-worker of mine at CIQ. Um And so uh Damon, I I won't I won't I won't steal your thunder. Why don't you give us a little bit of an introduction? Yeah. Uh it's great to be here. You know, as Eric said, my name is Damon Knight. I am a senior principal automation engineer at CIQ. And unofficially, I'm I'm kind of the in-house AI guy for for all things. So tell me automation engineer, that sounds like that sounds like fun. What what do you get to do with that? Uh you know, it's it's a lot of fun putting together an operating system like like we do.
But it's very difficult. There's a whole lot of steps involved and as we come out with more and more variants and we want to push releases out faster and faster, you know, it's it's impossible for the poor humans to to hit all of the buttons fast enough and keep everything up and running and and stable and properly tested. So I spend a whole lot of time making sure that us engineers get to sleep. As I get older, I found that sleep is actually important. So, so that's that's always a good thing. Um so, I I always ask my guests, you know, who are you and what do you do?
But, uh the the question that everyone seems to enjoy the most is what do you do for fun? When you're when you're not automating tasks away, working for working for a a for-profit company, what what do you do on your own time? Oh, you know, I'm I'm a nerd through and through. When I'm not writing code at work, I'm frequently writing my own code. But, on top of that, I'm actually a an amateur astrophotographer. I actually live out in the middle of the Mojave Desert. I've got a bunch of giant telescopes in my backyard, and I spend a lot of time taking pictures of the stars.
View full transcriptHide full transcript
Oh, that's so cool. As as a sci-fi geek, we'll we'll have to talk about that. That's that's really awesome. So, with that in mind, I did mention that this is the first officially sponsored podcast episode of the IT Guy Show. Back in my days as a pseudo show podcast host, we we had uh had sponsors basically from day one, thanks to Destination Linux. The IT Guy Show I've I've kept sponsor-free for now, but I'm kind of doing I'm kind of trying to highlight some really cool movers and shakers in the industry. And so, CIQ, the company that both Damon and I currently work for is is kind of the same.
So, or is is kind of a special case just because I've really all of you know, I'm I'm I love technology. I love cool ways of doing things. And so, we're we're going to talk about one of the things that CIQ just released a few weeks ago. But, just just disclaimer, Damon and I both work for CIQ, and yes, this is a sponsored episode of the IT Guy Show. This will be your informal commercial for the episode. The rest of the content will be will be kind of will will just kind of be the the typical format that you all have come to love and and download.
So, with that said, uh if you're interested in sponsoring an episode of the IT Guy Show, please reach out contact@itguyerik.com or podcast@it itguyerik.com would probably be a probably be the more direct route. With that said, let's get back into it. So, Damon, you said you you became kind of the the de facto IT IT AI guy here at CIQ. So, what what don't we start there? What's what what are what are your some what are some of your responsibilities as the AI guy?
Uh yeah, you know, a lot of my time is spent and has been spent on building out AI benchmarking systems and working on performance tuning and you know, helping to kind of teach you know, internally you know, what what best practices are for using AI and working with AI and and all of that and just kind of kind of help push push things in the right direction to to make it a a less painful experience for everybody involved. Yeah, that makes sense and the the more I realize the the more I work with AI both locally and commercially hosted options, it's it's amazing how just changing a few words in your prompt can go from wow, this output is utter garbage.
This is the worst tool I've ever used to oh my gosh, I never would have thought of it that way. Yeah, there's a bit of a joke that you know, AI is is an intern, right? It's sometimes it blows you away with with the the ideas and and the the cool stuff that you get out of it and sometimes you just shake your head and you go man, I what? Huh? What what were you thinking? Yeah, my my favorite is the comparison. It's like here's my old draft. Here's the new draft. What what works better? What doesn't? And it comes back and goes it looks exactly the same." It's like, I just spent the last hour rewriting the entire body of this.
This is not the same. But, okay, you go home, you're drunk. But, that that kind of that kind of leads to a more broad question. And and I I like to do this with with guests. I like to kind of start at the the 50,000-ft view and kind of zoom in on on kind of the specifics. So, from from your I mean, you've you've been in automation, you've been on the Linux side, uh and now the AI side. So, I'm really curious about your particular uh opinions about the state of AI infrastructure, particularly from an infrastructure perspective. Sure. Um so, you know, I I have some bias here.
I'll be honest. Uh and and it's not actually because of who I work for. Uh it's because I have a a ridiculous lab in my house uh that that I like to do testing and tinkering and all of that on, right? I've got I've I've got a a server rack in a shed in the backyard, uh as everyone does. Um and so, I I like to play with AI as as a user, right? Not not just as, you know, for my job and building AI things, but I like using it. Um and since I have my own infrastructure, I spend a lot of time dealing with my own infrastructure.
It's you know, the AI is amazing and wonderful and great. Um infrastructure hasn't really kept up. Um you know, I can't tell you how many times I've had to completely blow away servers and rebuild from scratch because I got lost in dependency mazes and you know, uh just winding up really frustrated and spending more time trying to figure out how to make it the AI system work than using it. Right. And so, I'm I'm I promised you that we we often go off script on the show and we we've done so already. So staying true to my roots here. I'm curious. How how is that balance?
Because one of the things I want to try and do around content around my own home lab this year is I want to become more reliant on my home lab. Right now it's running Plex and Minecraft and and a couple of other miscellaneous things that I don't use that often. And then in fact in in the previous episode I was talking about how how tired I am of all these subscription services. I'm tired of paying out $20 here and $50 there. I mean streaming itself is can we can we agree streaming has gotten worse than cable ever was and there's still nothing on to watch unless you're watching Star Trek Strange New Worlds.
But it's it's like I I want to bring more and more of that in house. And and so I'm I'm kind of curious as as I've been playing with this idea and I've been asking around the industry and and my my own network of of nerdy friends, what what is that balance like for you? I I don't want to become the at home systems administrator. There's a reason I went into sales and marketing was so I didn't have to be a sysadmin anymore full time. So what what does that balance look like for you? Well, I mean for me there isn't really a balance, right? This is this is what I do for a living.
I do it because I love it. So I'm you know a little more willing to eat some of the pain than some folks are. But you know, I'm a lunatic. I I run my own I run my own DNS. Like I own domain names. Like I'm not using route 53's name servers. I'm running my own name servers in my house. Uh for >> Okay, you've just you've you're you're on a different level here, sir. I'm I'm not I'm not a sane person. Um but actually honestly the the reality is actually I don't spend a lot of time on care and feeding and upkeep. Um getting it working the first time is where you spend most of the effort.
But once it's there, as long as you're being good about you know, backups and all of the standard best practices so that you're not real sad when the inevitable Yeah, because it will come, uh then then you're fine. But you know, I've got a petabyte of storage. Like wonderful it's wonderful astrophotography, man. It it takes up a lot of space. >> That's fair. But that also means there are no backups. Right? Where where am I backing up a petabyte of astrophotography? I'm I'm not. Uh so, you know, I have to put on my sysadmin hat and go, "Okay, cool.
Understand what I care about, uh you know, figure out what my processes are to to keep the the important things safe and accept the fact that yeah, I'm self-hosting and being in the in the self-hosting camp means there are risks and dangers." So, if anyone from Backblaze is listening to this, I think Damon would like to discuss your petabyte backup plan. Maybe Backblaze will sponsor an episode now. So, we're we're we're before I I you know, I'm the host and I'm usually the one to derail my guests from the agreed-upon conversation. So, getting back to kind of the state of AI infrastructure. Um kind of talked about what it's like kind of from your own perspective, but what about from the industry level?
What what's kind of the average organization dealing with as far as GPUs and AI models? Yeah, you know, the the AI market's in in a weird space, right? There's not enough hardware for anybody. Uh you know, resources are strained. So, a lot of companies are just honestly taking whatever they can get wherever they can get it. Um a lot, you know, tends to be in the clouds, right? The hyperscalers, you know, grabbing, you know, uh AWS, you know, hosts or or Azure hosts or whatever that have those big GPUs, um because it's it's, you know, your your options are limited. You can do that. You can spend hundreds of thousands of dollars buying servers and and GPUs, but otherwise, you're you're stuck with the old gear.
A lot of that, too, right? There's There's folks who are running, you know, T4s still in production today. And, you know, if Heck, go to Azure. They'll sell you A10s all day long. Those are There's nothing wrong with them. They're a great GPU, but they're they're getting long in the tooth these days. Um So, like all these companies have really mixed environments. Stuff is spread all over the place. They don't know what they're going to have. Um and and who even knows what operating systems they're going to try to Right. It's the the driver stacks are different. The install processes are different. Uh it's it's it's a little bit of a mess.
Uh it's And and, you know, if you work in IT or, you know, you're you're one of those admins trying to, you know, maintain your your Terraform scripts or or whatever to to keep everything sane and right is just It's it's not a good time. Hm. Yeah, I mean, I've I need to replace the USB-C cable for my webcam and uh replace my Nvidia uh gosh, it's it's consumer grade. It's like 7, 8 years old. I I forget what the model number is off the top of my head. But, about all it does anymore is just Plex transcode.
And I I tried to throw an a local LLM at it and I swear the thing called me on my phone and just said, "Please turn this off." So, I can I can only imagine what it's like There there was a story that leaked out of Nvidia I think last year even may maybe around super compute last year where a lot of the Nvidia sales folks are just kind of sitting back cuz they've already sold everything for like the next 18 months. That's how far back ordered some of these GPU chip providers are and there's there's just not a solution in sight and these AI models are getting bigger and more complex.
So the problem is getting worse not better. Oh absolutely. The the cards that were cutting edge, you know, three years ago are unusable and almost outright today in a in a lot of cases. That's That's crazy. So let me let me ask you another off the cuff question here. What what What do you think the solution is? Is Is the AI bubble going to pop or is there some new hardware platform that's going to come to be? CPUs get better at doing AI models? What What What does Damon think? That's that's a really tough one, right? I you know, it's it's obvious that the current approach won't scale for forever.
Um you know, fundamentally there's got to be someone somewhere has to make a breakthrough and then they probably already have and we just haven't heard yet. But right someone has to make a breakthrough that really brings the the availability and the accessibility down, right? If that's you know, getting it so that older hardware can Oh cool, we don't actually need so many tensor cores. You know, we can we found a way to do this in floating point math or in some other way, right? Cool. Awesome. Or or maybe it's just okay, we found a thing that isn't, you know, DRAM heavy. Mhm. Please, right? If we can find a way to to not be so so memory bottlenecked, that would I'd save a fortune.
Right, no kidding. And like like you I want to run an LLM in my home lab. Like I joked not not really joking, but I've I've made the comment that I would love to take some of the some of the D&D modules and rule books that I have in electronic format, feed them to a local LLM, and build my own my own DM. Build be able to just We should talk. We should talk. Definitely, for sure. Um So, yeah, I I don't have a I don't have an answer either. And you know, it's it's prices are going to continue to go up, and it's affecting RAM prices now, too.
And of course, the the global economic and political climate isn't helping things either. And there's just I don't know, maybe maybe this AI bubble pops. I I remember a decade ago when blockchain was going to solve every problem out there. And that that kind of died off. I don't think I don't think AI will die off quite as much. I think I think it's here to stay, but I don't think it'll look like it does today when it finally solidifies. I think the honestly, I think the the big shift is going to be um you know, if you've looked at those unified memory architecture systems, right?
The Apple systems, uh those those AMD systems with unified memory architecture, where you have a shared memory bus. Um I think those are going to make as as those get better and faster, um that's going to really open it up a lot. Like, I have one of those mini PCs um sitting, you know, in the other room. I'm looking at it right now. Um and you know, it's only got 128 gigs of RAM, and and none of that is dedicated VRAM, because it is unified memory. But, I can run Quake 3 Quake Arena next and have it be completely usable. Like, a little bit slow, but no worse than Claude when Claude is having a slow day.
Oh, and does he? But, right? Like, that's Okay, you know what? To me, that's holy cow. You know, for for under three grand, I can now have usable local LLM. That Wow, like that hasn't been a thing yet. Uh Uh like this is this is I think the big shift. And so I write the even if the the bubble were were to pop, right? Like AI coding, AI assisted development, AI workloads aren't going away. But they're going to shift where and how they run and it's you know, we have the the the the moon may no longer be be the target and you know, what open AI just shut down Sora cuz video AI generation is way too expensive even for that.
So maybe maybe some of those things start to fall off. Yes, yeah. Yeah, and I like that. That's that's a good pragmatic view of uh And and I think it'll happen a lot quicker than a lot of people are projecting. I like some of the model breakthroughs that people were expecting around 2030 or so, I I think we'll see in the next couple of years. Uh it's just it's accelerating so fast and as as someone who was stacking who started his career stacking literal servers and like I remember when VMware announced vMotion for the first time, the ability to move a virtual machine from one physical host to another without even dropping a packet.
When I started, that was revolutionary technology. That was going that was going to be the end of systems administrators as we know it and here we are, you know, in 2026 and and we're we're talking about building AIs that can build other AIs that can manage infrastructure. I mean, granted, this is the same system that couldn't spell strawberry a year ago, but you know, we're we're we'll get there. Honestly though, right? Like that I love that though, right? That to me is a wonderful example, right? If you say, "Hey, you know, where was AI literally a year ago?" Yeah, how many hours are in strawberry? I don't know.
Uh And and then you look at it today where you know, Anthropic is using Claude code to write Claude code and like co-work and and major features. Uh you know, they're they're dog fooding real hard and launching production quality businesses, right? That's that's not Granted, some of some of the other enterprise Linux companies have released their version of of what we're going to talk about today already. Uh, but being here and getting to watch it come out of beta and be public release, we're of course talking about RLC Pro AI. That is Rocky Linux from CIQ Pro artificial intelligence version.
We'll talk about what that is here in a second, but uh a lot of the different a lot of different Linux providers have released their own versions of this, but one of what what's what the industry has realized is particularly with the hardware limitations, particularly with hardware availability, we've come to realize something that hasn't been a problem since since the days that we left mainframe for consumer not not consumer grade um what's the word I'm looking for? Um We we left mainframes for for for Dell systems that are that are uh built out of components that we can we can build multiple of the same and um gosh, the the word's escaping me.
I I speak for a living and, you know, sometimes English is hard, but um you know, we we went to this this this space where you can buy a thousand Dell servers, HP servers, or whatever, identical, and and just throw them in a rack, throw them in a data center, power them up, install the operating system, and go. This this is what we call a general purpose Linux operating system. And if you go to anywhere, Rocky Linux, CentOS Stream, uh Ubuntu, if we go to any of these places and download an ISO, that's predominantly what's for the last 20 30 years what you get is a general purpose Linux install.
And then you you kind of build it to whatever spec you need. You install databases, or you add extra memory, uh and um you run some graphically intensive web applica- web application. But, I think we've kind of hit a wall here between hardware and cost of ownership and just the sheer complexity of computations with with some of these AI models in particular where the general purpose Linux install just doesn't cut it anymore. Yeah, I think I think you're really right about that. Um you know, we we talked about it earlier, but yet it's it's really common to see folks spend more time on the infrastructure than on the actual application or or whatever it is they're trying to try it it's and that's when you know things are upside down, right?
That's that's not how this is supposed to go. Wait, it it's not? I I know. I know. We've we've been doing it apparently we've been just been doing it wrong for the last decade or so. Well, I'm I'm reminded from time to time by a by friends of mine that not everyone installs Linux just to install Linux like the the sysadmin in me does. I used to distro hop back in the day, but apparently most people install Linux virtual machines or servers for specific workloads. So, that I don't know if if that was if that was news for you, Damon, but it was it's a shock to me.
Yeah, yeah, I you know, big very big surprise. So, what what is it about these AI workloads that's different from just a typical server workload? Well, I mean obviously right you're you're spending a lot of time in GPU space. You you're not CPU bottlenecked almost ever. If if you are something has gone wrong. But, you're spending a whole lot of time hopefully GPU bottlenecked or memory bottlenecked. But, like suddenly you're you're worrying about you know, bus speeds and things like that where again general purpose computing you don't care. You have memory and CPU and as long as you've got headroom, you're happy. Um and now you've got other devices that that you have to worry about and you know, the those devices are big monolithic blocks of of importance to you, right?
Um you especially if you're coming from the consumer space, you know, if you're dealing with a the you know, a a 1080 or 3090 or 5090 or or whatever, right? You if you want to run AI workloads, congratulations, that's what your GPU is doing. Uh and you're going to have one model loaded up and you're going to be doing like one thing. In you you head over to data center world and you've got technologies like MIG, you can split GPUs into smaller GPUs so that you can run multiple models and multiple kinds of things at the same time on slices of the same hardware. Uh but right, there's there's a whole lot of complication that comes with that.
There's all this, you know, software from Nvidia and you need special compilers and you know, what what libraries are are you know, is and it's not just what libraries does your hardware support, but it's what libraries was your software written to utilize. It gets real complicated real fast. Uh and you know, Nvidia likes to deprecate drivers and and support for hardware and so it it's you know, it's kind of a crap shoot sometimes. And that's that's a really good point, especially for anyone who's been around Linux for any length of time and has ever tried to install Nvidia uh GPU graphics drivers.
Um so any anyone out there is probably nodding their head and and grinding their teeth a little bit about trying to get Nvidia uh graphics installed properly on Linux and then inevitably someone forgets to recompile the kernel with with the right modules and and it's it's kind of a mess, but it's even more so when you talk about AI workloads. We we get into things like like CUDA. So could you could you explain to folks that aren't necessarily sure what CUDA is or maybe have heard the term, what is CUDA and and how does that impact your your operating system? Yeah, so CUDA is kind of all of the the magic software that Nvidia writes to optimize their GPUs, right?
It makes the software work the best. It what enables all of their cool features. If you have multiple GPUs, right? That's that's where the some of the magic lives. Yeah, for for NVLink and things like that. But right, the the CUDA stuff is what makes the the tensor magic and then the cards run fast. Um but that's also, you know, while that's a wonderful thing and it's why Nvidia kind of dominates the the market so heavily because CUDA is really good at what it does. It is a big piece of software. It's a complicated piece of software. If you frequently requires its own compiler to to write things for and there isn't a nice friendly way to install it, which is kind of interesting.
You know, at least in Linux. You know, in in if you're a Windows user for for some strange reason. I'm kidding. I'm kidding. But but you know, right? The the the process is a little more streamlined and a little easier, right? You can just you just go to nvidia.com and you download an executable and you run it. In the Linux world, it just doesn't work that way, right? It the CUDA versions depend on your underlying OS layer and what drivers you've got and what hardware you've got and all of that. Um so it's it's exciting. So do you or someone you've you've worked with have a a horror story around trying to get Nvidia drivers and CUDA set up and working?
Oh, so many. So many. >> The audience loves a good disaster story. We we all seem to thrive on others' pain. Well, you know, it's uh So I mean, I've got I've got my home lab uh and in the days before uh RLC AI, uh if if you wanted to run AI stuff in general, you were probably going to Ubuntu. Um that's that's just kind of, you know, the community is there and then so there there tends to be the space where you default to. Uh and who boy. Uh you would think Ubuntu being one of the friendliest Linux distros out there, installing CUDA would be easy.
But no, right? You you can't just apt get install CUDA. Yeah, you would think maybe you could uh because that's the the way most Ubuntu things work, but no, no, no, you have to like go get the repo, you have to get the GPG keys, right? There's there's all of these steps. And you have to also get as I mentioned before, you have to get the right CUDA version for your OS. Uh and uh Ubuntu has has interesting numbering systems, right? Uh 24 was was the the LTS release for forever. Um 25 was the the new hotness for a while. I think 26's out or coming out and I think that'll probably be new LTS again.
So we are releasing on uh March 31st and so 26.04 should be out in the next few weeks. Cool. All right, so pretty pretty close. Um but so right, there's if if you look for CUDA drivers, there are Ubuntu 24 CUDA drivers, and there are Ubuntu 26 CUDA drivers and packages. They don't exist for 25. Uh that doesn't You You can still do it, but you have to know that you have to install the 24 CUDA library, uh which is which is maybe not obvious. Uh I I didn't find it obvious, uh but yeah, I I spent a whole lot of time trying to figure out why I couldn't find CUDA drivers for an OS that clearly CUDA worked on.
Uh and then when I finally got that working, I I discovered that um the Torch is a thing, right, for doing AI workloads. So, you're you're dealing with with Torch, um but you have to compile Torch for CUDA, and you have to compile Torch for the right CUDA, and if you do any of those things wrong, you're you're doing all of your inferencing and LLM work on your CPU, and may or may not know it until you wonder why you're getting half a token a second. Uh which which is exactly what happened to me.
After I finally went through all of that nightmare to figure out how to get the the the 25 needs 24 CUDA drivers thing, uh I I wound up with the wrong Torch compiled CUDA, uh and so I was I was really confused why my A30 was really under uh and I'm not not generating tokens at the rate that I expected. It's Oh, oh, no, I'm still doing CPU generation because my the CUDA wasn't Torch properly, and Oh, I I feel that pain.
And And the one thing you you forgot to mention was that like I I've I've found this to be pretty much universally true that it is that if you're a systems administrator, doesn't matter how big your environment is, You'll inevitably end up with Ubuntu and Enterprise Linux distributions, whether that's Rail, Rocky, Alma, CentOS Linux 7, which is still out there in the hundreds of thousands of nodes that is just like Oh, great. So, which which version do I use on this? And the thing and the stuff and that's when you break out the the bourbon and you just know I want to be up until midnight trying to work on this.
I you know, honestly, I um just last week spun up a Ubuntu box to to do some AI stuff on. I was recording some demos. And immediately was, "Oh, man. Okay, hang on. I have to go Google and like start writing down instruction sets ahead of time so I can remember how to get this up and running because it is long and complicated enough that yeah, it kind of requires instructions. Uh Oh, goodness. So, we we've hinted at it a little bit just a few weeks ago and the the release will be in the the show notes. But a few weeks ago, CIQ actually released RLC Pro AI.
And so, the the the quick overview without getting with without putting on my marketing hat, the quick overview is that Enterprise Linux is a well-established platform to build off of. And I gave this talk so many times at at Red Hat that I can still do it in my sleep and do occasionally. Uh so, you you start with the upstream kernel and then from there the Fedora community pulls in and they build out a new release every 9 months or so and then every so often CentOS Stream will pull from Fedora and create the next major version of CentOS Stream and then every 3 years Red Hat will build Red Hat Enterprise Linux off of what it equates to kind of a snapshot of CentOS Stream and that's what becomes the next major version of rel.
Now, that's that's kind of the top half of of that uh tree. The next piece is then community Rocky Linux will actually rebuild from the rel binaries. Um and that's how we get community Rocky Linux and uh and then from there CIQ as a corporate entity very very separate um but kind of intertwined intertwined can separate but intertwined does does that work? We'll go with it. Uh but then CIQ pulls from community Rocky Linux and does some enterprise things to community Rocky and builds RLC uh plus and RLC Pro. If you're listening and you're looking for a CentOS Linux type experience, look at RLC plus and we'll we'll talk about a couple of reasons why that that'll come in handy here in a minute.
Again, I promise not to get too marketing-ish uh but and I'm trying to paint a picture, I promise. Stick with me. Um So, with RLC plus and RLC Pro, those are the commercial products that CIQ builds off of community Rocky Linux. Um and with that, uh there's ways that that CIQ changes the kernel, the platform, the user space a little bit from community Rocky Linux. Um so, while everything's still binary compatible, it's not identical. Those are very two very important terms to kind of keep in mind. It's it's compatible but not identical. Especially when you get into some of these RLC Pro variants. There's a hardened variant, so if you work in the federal space, you might look into RLC Pro hardened.
But today, what we're talking about what CIQ just released a few weeks ago is RLC Pro AI. And what this is is we we talked about general purpose Linux. But what this is is a purpose-built Linux. I don't want to say distribution because people kind of freak out when you talk about distros, but a variant of community Rocky Linux built on RLC Pro. RLC Pro AI has its own specific version. I'm trying to choose words carefully because definitions get overcharged. But, there's the CIQ Linux kernel or the CLK kernel that's used in RLC Pro AI. There's some changes to user space, the libraries that are included.
So, it's purpose-built. RLC Pro AI is purpose-built to run local LLMs and do inference, do training, do kind of your typical chat setup. It's It is specifically built for this. And so, RLC Pro AI is one of one of the first. I don't think we were the first, but one of the first Linux distributions to come out with an AI variant. And so, it's really, really cool because with with some benchmarks and and Damon, I'll let you talk to some of the numbers. Um, by purposely tuning your Linux kernel and your Linux distribution to an AI workload, you can actually see some pretty great performance benefits, right?
There's There's all kinds of neat stuff you can you can do just depending on the the specific tunings and and tweaks that you want to make. We We think we've come up with a a pretty decent starter set at least to to ship with. Yeah, and I mean, especially with with deployment issues, with all the engineering time that goes into this, with hardware availability, every little every fraction of a second you can take off processing time, every little bit of space that you can squeeze out of your hardware, just makes it so much easier. And so, I I mentioned RLC Plus and RLC Pro AI, there's additional benefit to this.
And that would be that all the drivers, all the PyTorch versions, everything, all the all the CUDA, all all the things that we were talking about are already pre-installed. So, CIQ and NVIDIA have worked together um and then CIQ and AMD have worked together to build variants on RLC Pro that already have the the drivers set up. So, all you have to do is put in your card, install the operating system, and away you go. And so, RLC Pro AI already has the drivers, already has CUDA, already has PyTorch preconfigured and tested um ready to go. And what's really cool is if you're on the Homelaber side or if you're on the uh community side, there's RLC Plus with NVIDIA.
There's RLC Plus with AMD. It's really it's really cool because it's already set up and it's it's validated. Uh and so, I I want to kind of poke at that for a second because it's one thing to just install the latest drivers, create an ISO, and go. But, the these are not pre-installed drivers, right? These are these are validated drivers. So, what's what's what's the What would you say the difference is? Well, so, you know, it's really easy to to say, "Hey, this is how things are supposed to work." Right? We can we can say, "Oh, yes, well, the the drivers are there. So, so, you boot them up, they should work.
Oh, CUDA is there. So, you know, you fire up your thing and it should work." Um but, you know, as I mentioned earlier, if you know, if you compiled torch without the right CUDA support, you're you only have CPU torch, things like that. And that's that's where we spend our time, right? We actually you know, we precompile things with the right flags. We actually tested things, right? We boot things up on you know, I've got an A30 here in my home lab. We did a lot of testing on on my portable blade. But, right? Like, we we fire it up in in the various clouds. We test it on you know, H100s and RTX boxes and and all of that.
And so, we validated that not not only do all of the pieces work, but they're also talking to each other properly. Again, you're you're not going to be recompiling torch because it was accidentally shipped on CPU only mode, right? Like no, no, no. We we handled that. We're we're shipping with not just torch actually. We I think we torch vision is part of our our pre-included and pre-configured libraries, right? Like the the things that everyone working in the AI space is going to need, they're just already there. And again, not just already there, but already there and talking to each other and working properly. Yeah, I mean, the the process that you're talking about of going through installing the drivers and making sure that everything's configured and and then ensuring that you're actually using the GPU versus the CPU.
I mean, all that the first time out can take a couple of days until like like you're saying, until you get the process documented, it can take a couple of days. Like just just a year or so ago, I remember it took me probably 3 days spread over like 10 10 days or so just to get my Nvidia GPU blacklisted at the Proxmox level so I could pass it through to my Plex virtual machine cuz I I like to I like to abuse myself with crazy use cases. But uh that that's just getting Plex to use an Nvidia GPU for transcoding. That's not that's not gigs and gigs and gigs of of storage of virtual memory for for an AI model.
So, I mean, the process just gets so complicated so quick. But with with RLC Pro AI, I can actually just plug in the ISO or deploy it on on one of the one of the major hypervisors and with I I I think step two is just download your preferred model. Step three, profit. Let's go. Yeah, it's one of the things that, you know, I I I like to mention to people as I as I think is a kind of a a neat uh selling point almost uh is you know I speaking to Ubuntu, you know, I've done this enough that I now have a I've written my own scripts that set everything up for me on Ubuntu and install all the right things and do all the right magic.
And And again, this is a script, so I'm not typing anything. Like this is best possible case scenario. It's a 13-minute process and may or may not involve one or more reboots depending on the specific hardware and if I have to blacklist new bug drivers. Right. But like but like that's literally that's me SSHing in and running the script that does everything. Best case is 13 minutes that I've managed. Oh, gosh. Fast, right? If you're if you're Googling, if you don't have a script ready to go, yeah, then you're 30 minutes, an hour, multiple hours is completely reasonable and to be expected. And you know, doing the exact same thing in RLC Pro AI, again, I don't I don't have to.
I I SSH in and it's it's ready, right? If I need specific extra libraries or models or whatever, I can download them. But if I don't and I just want to start Python scripts that utilize torch, I do not have to do anything to the system. I can just start writing code and go. It's It's so cool. I mean, it seems like such a little thing. It's like, okay, great. You You performance tuned it for for AI, but it's so much more and it makes such a difference, especially if you're I mean, some of these companies are buying GPUs by the hundreds and you know that they're deploying those the minute that those GPUs finally drop in in the in the shipping box.
And I don't know about you, but I don't want to spend 13 minutes per server trying to get this set up. And And like you said, that was automated, that was best case, and that was after what, two, three days minimum of probably trial and error. If you're trying to deploy a 50 node AI cluster, that's ain't got time for that. And honestly, too, right? You know, the these cards are so expensive. Those 10 minutes, I mean, I don't you're right. We we can we laugh, right? It's only 10 minutes. But like overall, you know, if you bought 100 or 1,000, if you're one of these hyperscalers and you're buying hundreds and hundreds of hundred thousand dollar GPUs, well, those 10 minutes really matter cuz those are 10 more minutes of revenue you could have been making selling services to customers or whatever.
And then again, it doesn't feel like a lot to say 10 minutes, but it adds up Well, I mean, every every startup out there seems to have some kind of AI angle on it right now. So, if you're a startup and your your survivability is dependent upon getting those GPUs in for starters, once you get them in, assuming your your VC funding holds until you can get them in, then yeah, like you said, I mean, every minute is a big deal. You're you're talking about potential revenue losses of incredible numbers that, you know, I could one one company goes under and I could retire based on off of what they lost.
Um and if if you're if you're like me and a buddy of yours had to to step step aside from from DMing a campaign and you suddenly find yourself DMing two D&D campaigns, you know, those those AI models are pretty important. No no one's going to feel sorry for me for DMing two D&D campaigns, but you know, I just threw that out there. So, what what what's next, Damon? What where's RLC Pro AI going from here? What's what's what's the what's the goal? What's the hope? Well, you know, the the goal of course is, you know, to to be the preferred platform for for everybody who wants to do AI stuff.
You know, that in the future that might mean uh you know, including more libraries that the the AI community tends to want to use, more software packages, things like that. We're still we're still playing around with uh specifics there, so I won't give too much away. Uh but so you know, so again, it's more of that just we want to help you hit the ground running faster stuff. Um but also we're continuing to spend time looking to see, "Hey, where where can we get performance improvements? Where can we get efficiency improvements?" Um you know, we uh outperform uh because of the specific way we've built and precompiled things in certain areas of AI utilization, right?
AI segmentation. Um I think in in particular, uh image segmentation, I should say. Um we tend to because of our tunings outperform uh compared to other OSes running similar software stacks and and the same process. Um So, you know, we're going to spend more time looking at things like that, seeing you know, what what knobs we can tune uh that make the the lives better of of everyone. Uh you know, I'm again, I'm I'm selfish because I do this myself, so I want to make my own life better. Yes. Uh and if it helps everyone else out, too, all all the better. Uh so, it's it's a a fortunate confluence of events, I suppose.
So, if if uh if there's a sysadmin or someone in my audience that is dealing with this right now, uh what what advice would you give them? Just kind of general advice, whether they're an AI engineer or a systems administrator, um what what would you what would you recommend? Uh you know, first off, uh please please check out the the products from CIQ and and uh RLC uh plus and pro uh with you know, kind of those preloaded drivers and software packages. Um Do honestly, truly, I'm not just saying that as an employee, they they do help. Uh they make life a lot nicer and easier going through that process.
I do this a lot. Uh I promise. Um But But on top of that, right? Um you know, if whether you're using RLC or or anyone else's offerings here, um do your research. Uh do do that Googling. Figure it out ahead of time if you possibly can. There is nothing worse than, you know, spinning your wheels and running around in circles because, you know, you you forgot a flag at install time or or whatever. Uh right? We've We've all been there. Um You know, that's that's really the the key. And if you don't, I'll be honest, throw Claude code on that sucker and ask it to fix your system for you.
It's really good at that and not enough people take advantage of that ability. You know, honestly, I I just in the past 2 weeks have retired my ChatGPT subscription and have moved over almost completely to Claude. Um I haven't deleted my ChatGPT account because every now and then these these big corporate models will have a day where it's just like, "Nope, I don't feel like I don't feel like doing the AI thing today." And you ask it the same question three different ways and get no response. You Then it's really fun to tell one model that the other model is struggling. That seems to help grease the wheels a little bit.
Like, "Oh, so ChatGPT can't figure out this this blacklisting uh Nvidia GPU issue for you? Here, this will help." You know, it's it's funny because obviously I'm a I'm a big Claude user. We use Claude a lot internally at CIQ. I'm I'm a huge fan. Um but it does go down. It does break. And that local LLM box with its unified memory architecture, I can run open code against it. And so if Claude goes down, I mean, okay, it's all kind of a pain to switch, you know, switch agents and models and all that mainstream, but I am at a position where I can just keep going.
And And oh, by the way, that box is running RLC Pro AI for what it's worth. See, I'd I'd like to get to the point where I'm the opposite. I'd love to be able to default to my local LLM first and then and then go out to a corporate model. Um Six six months to a year and I think maybe we'll be there. Yeah, it's it's not easy and my my my interests and my my projects are so diverse between marketing projects, between planning projects. I've got um There's times where it's almost almost like I've got a project that's all I should just rename it to just brain dump.
It's like every prompt is like just paragraphs of this is the 16 things that are in my brain. I'm just going to dump it in here. Tell me tell me how to, you know, un tell me how to untangle this this spaghetti so I can get something done. Um and then like D&D and then like fact look up and and that kind of thing. And and it's just so many different things to to go into. And and then you've got a tool like In-A-In that I want to get into that you can actually have it try your local LLM first and if you don't get the right answer, it'll go out and use use your tokens with with one of your paid subscription models and it can do all this stuff.
And I I just I need to take a week off and do nothing but home lab projects and then just live stream the whole thing. What's what's also great fun if you're playing with a local LLM and you have a subscription is go ask your subscription, "Hey, I have a local LLM. Write me some tests to run against." I I So, I did this. I had Claude write me tests and then I fed them into my thing and I took the results and I brought them back to Claude and I said, "Hey Claude, this is what it said. What do you think? How does it How does it measure up?
All right, was this was this good? What would you have done better or what would you have done differently rather, right? And and >> Is is is that where it comes back and like rates it based on like like an ABC grade system and like Uh so it it didn't uh for me at least it didn't do that. It mostly just kind of like directly addressed, "Oh hey, yeah, I think this is roughly the same response I would have given." or it would go "What the heck happened here?" Uh definitely have some what the heck happened here uh when I was testing various models. Uh Uh you know, the local >> The local arms, they they're great, they're wonderful.
Uh but, you know, using them with tools and other things like that, right? They'll they'll frequently put uh you know, context or formatting in the middle of like answers or or things like that and you know, again, these are things that Quad Anthropic and and Open AI spend untold money uh you know, paper wallpapering over and making pretty and so you don't see it. When you're doing your own LLM development or using these local models, for the most part, that same level of care isn't there. So it's up to either you to either have to rely on them to to get their act together or you have to, you know, write your own.
I've written my own code to to fix uh you know, I I wrote a I wrote a proxy for uh Qwen 3 uh pre-mixed uh because it couldn't work with with tools. They would break tools. I wrote a proxy that would just intercept all of its tool requests and fix them and rewrite them correctly and then send them on. >> Wow. That's it's a crazy world we live in. And you know, for a few years I was like, "Has technology just gotten dull? Is there is there just nothing new to talk about?" And then like this this whole the the last few years it it started out like, "Oh, this is just a fad.
This will go away." Uh Um but there's just so many new and different exciting things you can do now. Absolutely. I I was the biggest AI skeptic in the world. Uh and and then uh someone convinced me to try cursor uh about a year ago. And my life has just never been the same. I I drank the Kool-Aid. I'm I'm I'm all in. Uh it's it's amazing. Honestly, it it for me it it kind of rekindled my love of engineering and building stuff. Mhm. I love that. It's so much easier. I I don't have to spend days and days and days running down dependencies to to try out an experiment or to try out a cool idea.
I can go, "Hey, let's get something sketched out. Let's get something that I can start banging against and playing with and seeing what ideas work and what don't." Mhm. in minutes instead of days or weeks or I give up after 3 months of not being able to Well, that's awesome. Well, uh any any last thoughts you want to share? Something we've talked about or something we didn't talk about? You know, uh I I think I would mostly just say, you know, everybody out there uh keep your eyes open. The The industry is moving fast. Um CQ is moving fast. Uh I think there's going to be really awesome stuff honestly just kind of all over the industry um over over this year.
I think this is going to be a really really exciting year for AI, especially after seeing what what happened last year. Um and I I hope as a long-time sysadmin systems engineer uh that, you know, the the pain of of AI is is a thing that that we're hopefully going to going to solve or or at least make a whole lot better. Um Cuz I again, I care a whole lot about it myself. For sure. For sure. Well, that's awesome, Damon. I really appreciate you coming on the show and uh and >> It's been an absolute pleasure. Thank you so much for for having me on.
It's been great chatting with you. Yeah, and and uh of course thank you to this episode sponsor CIQ aka Control IQ. Uh really excited to to be working with uh with CIQ on some of these technologies. It's it's I I've learned a lot about high performance computing in the last few months um and getting to see how that's kind of fueled this this move towards AI and purpose-built uh Linux variants has been really cool to see. Um so definitely go to ciq.com to learn more about RLC Pro AI. And if you are a hobbyist and a home lab user um check out RLC Plus. Uh they come with a couple of different variants with more in the pipe.
Um the the biggest one we probably talked about was RLC Plus Nvidia where the entire Nvidia graphics stack is already pre-installed on top of a Rocky Linux base. So it's uh really makes uh really makes life easy. Uh so that brings us to the close of episode 20 of the IT Guy Show. Uh starting to There was There was a long pause there with uh with the holidays and with the move and with the job switch, but uh glad to be back on a semi-regular uh cadence. And if uh if you want more uh Eric the IT Guy, head over to the Fedora podcast. Next week we'll be talking about Fedora Flock.
That conference is coming up before we know it. Uh I'll be working with Noah Chelliah and the uh the the Fedora podcast team to be talking about Fedora Flock, the on-site uh developer conference that the Fedora hosts every year. Uh so check out the Fedora podcast and make sure to like and subscribe this content because I've I might just have an expert home labber coming on the show to talk about his build uh and uh and maybe give me some advice and hopefully all of you some advice before I really start trying to figure out uh what I'm going to do with that PowerEdge uh R730.
No, it's an R720. Sorry, R720. It's even older and really really long in the tooth. It It needs to be laid to rest, but uh lots of fun content coming up. And if you would be interested in sponsoring a future episode of the IT Guy Show or sponsoring the show in general, I would love to hear from you. Reach out at podcast@itguyeric.com. On behalf of my guest today from CIQ, Damon Knight, I'm Eric the IT Guy Hendricks. Really appreciate you all tuning in and hope to see you back here real soon. Thank you much.
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
9
Enterprise products
Spanning the kernel to the orchestrator
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.
