How to deploy your own LLM and take it to production with Fuzzball webinar poster

How to deploy your own LLM and take it to production with Fuzzball

Watch Now

Running your own language model, the list of prerequisites gets long fast: compute provisioning, GPU allocation, model downloads, service wiring, storage configuration, and authentication, with no guarantee it stays running once you get there. Many teams look at that list and reach for a commercial AI service instead. The teams that don't spend months on infrastructure work before a single model reaches production.

Fuzzball removes that overhead by capturing AI model deployment as a reusable, templated workflow. In this webinar, we'll demonstrate two ways Fuzzball supports AI deployment:

  1. CIQ's own tutorial video generation workflow, running live in Fuzzball on Nvidia DGX hardware.
  2. A self-scaling LLM that grows from a single team's use to full production serving, without changing the underlying workflow.

Webinar Synopsis:

  • How CIQ uses Fuzzball in production to generate its own tutorial videos on NVIDIA DGX hardware

  • How to deploy your own LLM as a reusable Fuzzball workflow that scales up and down on demand

  • How sovereign AI stays on infrastructure you own, with your data inside your environment

  • A live view of two real AI workloads running in Fuzzball, including one CIQ runs in production

  • A step-by-step understanding of how to deploy and scale an LLM on infrastructure you control

  • Confidence that the same Fuzzball workflow extends across hardware and environments without a rebuild

Speakers:

  • Wolfgang Resch, Research Computing Engineer, CIQ

  • David Godlove, Technical Product Writer, CIQ

Transcript

Hi everybody. Welcome to the CIQ webinar. So today we're going to be talking about using fuzzball and using it in particular for AI and to be very specific using it for sovereign AI. So a lot of teams right now are using AI for development work and debugging, writing documentation, all kinds of things. Um and they don't want to be outsourcing all of that AI to a third-party anymore. And they want to be doing something called sovereign AI where you actually are hosting your own inference models locally and you are you know your data never leaves your own machines and it stays exactly where you put it.

So that's going to be the subject of what we're talking about today and some of the ways in which fuzzball enables you to be able to have sovereign AI running in your in your infrastructure. So we got two different guests here that we're going to be talking with today. One is Hideto Harada who is kind of a new intern working at our newest intern working here at CIQ. So how you doing Hideto? >> I'm pretty good. How are you? >> Pretty good. And then later on we're going to be bringing Wolfgang Resch in and he's going to be talking a little bit about the same topic as well.

We've got two really great demos that we're going to be sharing with you today. Um the first is something that Hideto's been working on since he started with us. Um he's going to be running so it's a really cool demo. It's going to be running on an Nvidia DGX and it's going to be it's really cool because it's talking about how CIQ is using fuzzball ourselves internally to actually solve a real problem that we've got and produce some stuff that that is important for us. I'll talk more about that in a minute. That's a really exciting demo.

And then the second demo that Wolfgang is going to do is going to talk about how to run an inference model and then automatically scale it up and scale it down using Dynamo so that, you know, if you are hosting this for a large group of people to all be hitting the model at the same time, you're not going to run out of resources. It's going to go ahead and scale up for you. And you know, you're going to be able to host that for a large group of people. Um so, feel free at any point in time to drop questions in chat and we're we're here to answer even though, you know, this this is actually pre-recorded, but we're we're actually sticking around and we're here to answer your questions if you want to, you know, drop them in chat.

View full transcriptHide full transcript

Okay, so let me go ahead and talk a little bit about the first demo that we're going to do. Um so, we we had a kind of a problem we ran into a little bit of a problem here at CIQ where we wanted to produce a lot of videos about FuzzBall. Um so, we want those to be So, it's kind of funny because I'm going to be talking about training. And when we talk about AI and training, usually that means a very specific thing. We mean something totally different here. So, we want to create videos that are training videos. In other words, they're going to train people.

And we also wanted to create some promotional content for FuzzBall so that people know more about FuzzBall. And recording these videos is very labor-intensive. Um you know, usually if I'm going to do it, I set up a green screen and I, you know, put a bunch of stuff together. I use a different camera. Um and you know, we record a bunch of takes, we cut them together. It takes several days to record a single video. And and we want to put out a lot more videos a lot more quickly than that. Also, things are changing with FuzzBall quickly enough that we can't you know, take so long to record videos.

So, the idea came about, why don't we use AI to create these videos for us? And around the time that we had that idea, Hideto came on and joined the team. And he's got access to a G DGX. And so it was a natural fit. And so he went ahead and installed you know, fuzzball on the DGX. And he very quickly was able to create a template in our workflow catalog that allows you to quickly and easily give it a prompt and generate a video. And so that's that's kind of the demo that we're going to be doing today. It's talking about that. So Hideto, do you want to maybe take it away a little Do you want to say anything before we actually jump into the demo or you know, talk about that at all?

>> Uh sure. Yeah. What this workflow does basically just you give it a prompt plain English prompt. It plans the tutorial writing and like using an AI model using a a coin coder on on my GPU on my spark and it it creates a plan. It creates a script. And it drives a fuzzball web interface and it records that and it voices over it and it kind of walks you through a tutorial on fuzzball. And yeah, how to use it. Uh Yeah, that's that's about it. >> Cool. Let's let's dive in and actually look at it if you don't mind. >> Absolutely. Yeah, so this is what you first get when you open up fuzzball.

Uh it's it's a catalog it's a catalog Well >> Well, let's just back up just a minute. Sorry Hideto, sorry to jump in. Um so if for some people might be new to fuzzball all together, right? So so it'd be good to talk a little bit about what fuzzball is in this context. Um so Hideto has installed it on his DGX, but you can install fuzzball on AWS. You can install it locally on any you know, on a cluster. you can install it anywhere you like. And FuzzBall, um, you know, it it at its heart, it's a container orchestration platform. And it's a container orchestration platform really for high-performance computing and for AI workloads.

Um, here we're looking at the web UI. It's important to note that, uh, everything that we're looking at here in the in the web UI is all API-driven. And I think that, you know, we were talking about this a little bit before the session and Hidetoshi said he actually prefers the CLI. So, anything that you can do in the web UI here, you can actually do in the CLI as well. So, I think you were saying that you actually you actually drive this primarily using the CLI. But today we're going to be, you know, showing the web UI just to kind of, you know, make sure that people are up to speed and kind of kind of, you know, it's helpful for demos to kind of show the web UI.

Sorry to jump in and and and and take over there. Um >> No worries. Yeah, so this is the web UI. It's just the easy-to-use just simple version of FuzzBall. Uh, it comes with a free catalog of all of the fuzz files that, uh, run different scripts for different tools that we have. >> Yeah. Yeah, absolutely. I want to unpack that a little bit more, too. So, that we have this workflow catalog um, that has all these templates within it. And, you know, these templates are all pre-created for you by CIQ engineers and they come with FuzzBall. So, that the page that Hidetoshi was showing you previously has all these pre-configured workflows that'll run out of the box for you.

However, you're also able to make your own entries within the workflow catalog. And that's what Hidetoshi is showing you now is that he's actually created his own template, um, locally on the DGX and he can just reuse that over and over again. >> And yeah, there is also a copy button, so you can copy and edit your own version of my template and, uh, honestly put whatever uh, brand you want on it and have it run any tutorial for any company. Any companies UI that you would you would want to use. >> So what does it look So what does it look like to actually run this workflow, Adat?

>> Uh right now it's just setting your English for your plain English prompt. And then setting how many cores you want your AI use, your how many cores you want to use for your render. And uh the speed of how fast you want the narrator to talk. And you can turn a little I got a little mascot in there. He talks. He talks you through the tutorial. >> Does he? >> And we have a little thing to turn the intro on and off. >> That's >> That's awesome. >> These are just a few little parameters that you can set for the for fuzzball to work around.

You can set more or less depending on how much interaction you want. And after that you just click run. We already have one running right now. You can see it. It'll take you to this screen here. It'll show progress and you can click through and see your logs and events and see what's happening in your uh in your fuzzball. See see what see what you're working with. >> Now >> Yeah. So this is just the timeline of events that's happening. So first it creates or calls on the model to create the little script. It creates all the volumes that that stores the renders, the voice. And yeah, and that's it.

And then it sends it pulls some images and then it sends the little crawler after your your website. And it kind of clicks through and records that. And then puts kind of all edits it together in the render service and then it'll port it forward to a local host webpage where you can actually play through it and see what's going on. >> Now, that's fine. So, this is a training video. I mean, once again, we're using training in a very different sense than we would normally use it within the AI um sphere. But, this is a training video. It's It's worth pointing out that Hideto just created just now.

And, you know, this this video is is now you know, you we could put this up on the website and I mean, we check the accuracy of it and as long as it's it's good to go, then we can um use this to show people how to use foosball. So, it's a really really really really great thing that you can just put these together that quick. Um you know, it's cool that you can do all this within the workflow catalog and that makes it really easy for anybody to just be able to jump in and run these. Um Well, do you want us Do you want to show really quickly?

It looks like the the server started up. Do you want to show how you can actually connect to the server and interact with it in real time as well? >> Yeah, and usually you would have to like port it forward through the command line. But, it's nice enough to where it just puts a little button there and it uh it goes and kind of just opens it right up for you. And then, here you go. You can see it. >> That's really cool. Um So, it's really cool that you can do this uh you know, within the um we're we're kind of we're kind of um uh getting up to time here.

Um so, I'm going to be kind of wrapping it up fairly quickly. But, I was going to say it's really cool that you can do this within the workflow catalog. Um can you also show how you can actually open So, So, the workflow catalog, I should go back and say has templates that enable you to be able to have all those little buttons and prompts and everything that Hideto showed. So, that a new user can go in and just interact with a workflow and run it. But, you can also open this up in the workflow editor. So, if you go back to the actual running workflow, um and then uh actually go to the Well, there you that's actually the template.

So today to if you go back to the workflows and then click on the running workflow and then go open in in the editor in the upper right hand corner. So now we can actually get in and mess with the actual gamble file that was rendered by that template and so you can get really low level and you know create anything that you want here and you know change it around. It's really powerful that you're able to to do that and it's great that you got the flexibility to be able to give this to a you know kind of a new user or you can actually give this to an expert and they can do whatever they want with it as well.

So you know just to kind of like recap you know CIQ had a real problem that we really want to use sovereign AI with our own DGX and be able to create these these videos really rapidly and be able to iterate iterate on them quickly and it was so easy to do that we had today to come in and in a matter of you know what weeks you went from not having seen fuzzball at all to being able to install it bring it up create your own workflow run it and actually you know start solving a real problem that CIQ has and and dog fooding our product at the same time right?

So maybe now we could transition away from you know looking at the looking at the demo and just sort of talk a little bit more about you know what what what your experience was like using this and just kind of like you know discuss that a little bit more. So so how was it to be able to use fuzzball to be able to create this and just to start solving this problem? >> So yeah as you said I am a intern and I have really no previous experience with uh any software engineering or anything like that. So, EasyUse was actually surprisingly easy. You can Vibe code most of this stuff.

And with uh it's it's pretty straightforward. Like, you can kind of see everything that's going on. It gives you really good error error logs. So, when you do break something, you can fix it very quickly with the help of honestly a local model that that works. The the Quan model is very good for for Vibe coding this stuff. And yeah, it's uh Shoot, it's just uh uh it took me maybe a few days to get used to it. But after that, I was able to run every workflow through the CLI and see every log and uh I was able to get work done fast. >> We were >> quite a bit easy.

>> Sorry, we were talking a little bit um previously about uh this workflow and how Playwright is one of the places that's a little bit more kind of fiddly and brittle because, you know, Playwright is the is the piece of this workflow that's actually pointing to different things and clicking on buttons and interacting with the website as it creates the video. And we were talking about how um that needs some adjustment, but how FuzzBall actually enabled you to be able to do that quickly and easily because you don't have to worry anymore about tearing the workflow back down, allocating resources, spinning it back up, doing you know, having it reproducible.

All that stuff is taken care of of you is taken care of for you and you can just iterate, right? And just iterate very, very quickly and be able to get everything lined up. You want to talk a little bit more about that? >> Yeah, iteration is pretty easy. So, through the course of this, probably within a few days, I went through around 40 versions uh of this, just breaking it and fixing it, and breaking it and fixing it. And I could go back, cuz I have all of my templates saved, and I can see what broke and what got fixed, and what works and what doesn't.

And yeah, Playwright was kind of hard to work with, but getting up and running with it, like getting a base model that showed the screen click through tabs, was took probably 2 hours. The rest of the time I spent trying to map out the website. >> That's awesome. >> Button placement and all that. >> Cool. Well, thank you so much, um thank you so much today too, number one, uh you know, for joining CIQ and for doing all this great work. I mean, it's really appreciate It's really appreciated because like the work that you're doing is actually replacing a lot a lot, not all, but a lot of the stuff that I do, you know, and and a lot of the the other folks do, as far as creating videos and everything.

And also, thank you, um for kind of on short notice, uh jumping in and being a guest on this webinar. I really appreciate that. It's fun talking to you, man. >> Absolutely. Thank you. >> Yeah, cool. All right, excellent. So, um So, now I think that we want to transition a little bit, and we want to start talking to Wolfgang. So, the problem that that Wolfgang is trying to solve or did solve uh pretty quickly within Fuzzball, is the problem of Okay, so you you you've got an inference model that you've, you know, you've trained up, and it's ready to go, and you want to serve it out to a large group of people.

Fuzzball allows you to do that really easily because these services that run within Fuzzball, you're able to change the scope and make those either accessible, you know, just by yourself, or to everybody in your group, or to the whole world if you want to, as well. But, what do you do when those services start getting hit by lots and lots of people and it starts to overwhelm the resources that you've allocated? Well, that's where Dynamo comes in and auto scaling comes in. So, you're able to scale up and scale down these inference models in order to handle, you know, whatever the demand is that you've currently got coming in.

Um and so, he's going to show you an example of that. And once again, we're talking about sovereign AI. We're talking about hosting all this stuff on your own infrastructure, on your own machines. And that way, all this data, if you're serving the, you know, a model for your entire company, for instance, maybe you're serving a coding model that allows all of your engineers to be able to, you know, use AI to be able to debug code and create new code and stuff. Um what you're working on never goes out to a third party, right? It all stays local, it all stays on your machine, you don't have to worry about anybody else having your IP.

So, I'll hand it over to Wolfgang and he'll go ahead and do that demo and and talk about that for you. >> Hello. My name is Wolfgang Resch and I'm going to show you how you can create services in Fuzzball that scale automatically in response to load. I'm going to do this with two different workflows that I'm going to share with you. The first one will explain the basic mechanisms of how auto scaling works in Fuzzball. And in the second one, we will actually create a auto scaling VLLM based inference engine and connect it to local open code running in a Docker container. Okay. The first workflow we will start here from the Fuzzball workflow catalog.

This workflow has two services, a worker service that scales itself depending on how much CPU utilization it records. The second one presents a front end that load balances between the different replicas of the the worker job. Um there's only a few things we can change here. Uh let's say for example that we want between one and three replicas of the worker service. And I'm going to start running this. And while we're waiting for this to start, I'm going to open this in the editor to give you a basic idea of how the machinery is is wired up. First, the worker node. The worker uh service is basically uh runs a Python script.

Uh in in this example, this Python script uh every time in response to a to a request to one of its endpoints, will start a process that simulates a calculation and occupies one CPU core for 1 minute. Um this is a persistent service and it's it's defined generally, like all FastFlow services are, it runs code, it exposes an endpoint. Um one for checking whether it's healthy, one for actually submitting work. Um and the thing that makes this particular job auto scaling uh and more interesting than a general service is this tab over here. Um here we can, again, determine how many minimum replicas should exist of the worker service, how how many maximum.

For a service that scales itself, um you need at least one to start with. And then we can define triggers for scaling up and scaling down. Um let's look at the scale up trigger here. Um we're not running uh we're not running a particular uh server emitting metrics here, but FastFlow itself collects uh several basic metrics like, for example, total CPU usage in and CPU nanoseconds as a monotonically increasing counter. So, we can, for example, say if the rate of increase, meaning actual utilization, uh on average over a minute exceeds 1.5 in any container, start a new replica. And there's a cool down period that we define, meaning that if this trigger fires, another trigger can't fire for at least 30 seconds.

That pre- That prevents, for example, if if you have multiple replicas at scale uh reaching that that threshold, um they won't all fire um scale up events at the same time. Uh similarly, there's a scale down if load falls below 0.2 cores on a particular replica, it will get scaled down. And there's a cool down period of 60 seconds. So, that's that's the worker part that that makes this work. Um the front end presents an endpoint for the user to connect to, and it uh round robins requests per coming from the user to the different worker replicas. And the way it does that is by using a dynamic config.

Um it gets injected at some regular interval um an environment variable that um lists all available worker services. So, at the beginning, this will just be one worker node, and if there's a scale up event, it will show two. Uh and it writes that to a temporary file, and that code in the front end reads that every time a request comes in and picks one um one worker at random and submits work to it. So, that's the general idea. There's an auto scaling block that determines scale up and scale down triggers, and that's that lives in the worker here, and then there is the front end has a dynamic configuration that tells it how many workers are available.

That's what it looks like in the user interface, and if we look at the the actual uh textual um representation of this particular workflow, um we can see Let me see if I can make this bigger. Um so, in the worker, um this is the part that determines auto scaling. So, it's a straightforward translation of what we saw in the user interface. Uh we have a minimum of one replica, maximum of three, and then there's just the the formulas that determine scale up and scale down. Um I didn't mention this earlier, but these formulas are PromQL queries. And again, even without the job emitting any metrics, even though the jobs can emit metrics, um Pulsar itself creates several synthetic metrics, including CPU and memory usage.

So, that's that's that's the worker side of this. And then on the front end side of this, um we have the dynamic configuration. Again, this is what gets populated with the list of available worker nodes at a at a regular interval. Okay. And this is what presents the front end to the user, and it submits work to the worker service. So, going back here, um where our workflow started, uh and in this case, we have the front end that's already running, and we have two inactive uh two replicas that have not started, and one replica that has started, consistent with us requiring a minimum of one and a maximum of three replicas.

Uh this particular front end um exposes uh uh uh a um web interface we can connect to, and I'm going to put that in a split view here. And basically, all this service does is I can push this button, and it will send a request to a randomly selected replica um that is available, and that will then do a calculation for a minute. And we can see there's one idle replica, and three and two replicas that haven't started yet. So, now if I submit, let's say, five requests, because there's only one replica running, they will all get routed to that one replica. But, if we look at the events for this workflow, and let me make that a little bigger.

Um we should see within about a minute that a scale up has started. And the reason why this takes a minute is because we created a trigger that requires a load of greater than 1.5 over on average for a minute. So, it will take about a minute for the first replica um scale up event to happen. And then it'll take a second for that actually um being instantiated. So, we'll we'll wait just 1 second here. Oh, and here we are. So, you can see um that uh submitting several requests here uh triggered a scale up um of the workflow service and we have a second replica that's started up.

And now if we send requests, uh we can actually see that they get rep routed to different replicas. So, you can see that counter increasing uh in both of them. And that should actually eventually trigger trigger another scale up event, which will um instantiate replica number three. So, we'll wait just a second for that. And here we are. We just got a scale up event. And now we have all three replicas running. And if we send requests now, they will get load balanced between all three. And we'll let this finish, and then we'll see it automatically scale down um once these computations these synthetic uh computations have finished.

Okay, we're back after a little bit. Um and we can see that the first uh first scale down event has happened, and replica one was scaled down. And we're down to just two running replicas now. And in just another minute or so, we will see the second one scaling down. So, we'll be back right then. And here we are. We have that second scale down happening, and now we're down to just one running replica plus the front end that actually distributes the work between the replicas. So, you can see the same thing in the user interface that's presented by this replica. So, that's the general idea of um workflow services, and let me just recap um and and show you an illustration of how this um of the bits and pieces that go into making these work.

Okay, to recap um what makes an autoscaling service work in Fasb. First of all, you need a service that defines an autoscaler block. Uh in this case, this service scales itself because it always has one replica running, so it can detect when that replica gets busy and start a new one. Um the autoscaler block defines upper and lower bounds for the number of replicas, and rules look for instantiating new replicas, and rules for scaling down the number of replicas. Um in this case and then we also need this front end service. This front end service does a couple of things. It presents that user interface that I showed you to the user.

Uh and it submits work to the worker instances. Um the way it knows how many instances there are and how to reach them is through a bit of dynamic configuration. So, this service has a dynamic configuration block in its workflow definition, and this is how Fasb injects the the how many replicas there are and how to reach them. Okay. On to workflow number two. Uh as I mentioned, um this will be an autoscaling workflow that runs a VLMM-based inference engine as the scalable part of the workflow. Um and then it has a um a front end service, which allows the user to connect to the the load balanced replicas, and connect it to a local open code instance.

So, there's a couple of different there's a couple of things that are uh added to this workflow that we didn't have in the previous one. So, let me just go briefly over those. Um, first of all, uh here is the actual worker process in the in the fastfile.yaml file. Um, what you can see here is that the autoscaler uses a different kind of metric. Uh the autoscaler to decide whether or not a new replica should be started, looks at a metric exposed by the LLM. Uh and in this case, the metrics we're using is the number of pending requests. So, if there's more than if there's more one or more pending requests uh over uh over a certain period, then it will trigger a scale up event.

And if if nodes are idle, it will trigger a scale down event. And this is made possible because of this new block that we didn't have in the previous example, uh where we're basically telling fastfile that there is a Prometheus compatible uh scrapeable endpoint presented by this by this workload. Um, and that's this is the slash metrics endpoint presented by the LLM server. So, in the previous one, we just used a synthetic metric that fastfile collects for all services, which was CPU utilization. In this case, we're using a specific LLM based metric presented as a Prometheus scrapeable endpoint. And that's used to scale up and scale down um the number of replicas.

The next thing that's different, um this workflow has an annotation called fastfile.io/connect/the name of the service. In this case, this gives special instructions for how to connect to the front end service. And the front end service, like in the previous job, is the service that load balances between our different backends. And here what happens if we connect to this workflow from the command line, uh it will execute the script which will generate code. Uh it will use that to inject information about how to connect to our workflow, uh in particular credentials needed, uh what the name and the end the name of the endpoint is, and so on.

And it will then execute that local. And this in in our particular case, we're using this to basically obtain the credentials and configure a local open code instance running in Docker to connect to the load balance service to that to the auto scaling and load balance service in FaaSball. Okay. Let's go and get this done. Let's submit this workflow. Let me switch over to our web interface. Here it is. And here we see it's starting up. So, while we wait for it to download the model, start the front end and the replica, um we will be right back. And all our services have started up. Let's have a quick look at this just to make sure that it all looks healthy.

And yes, we can see that the worker running the inference model with VLLM has started up and it's ready to accept connections. So, go back to the command line and let's connect our open code session to this. And it needs one command line argument, which is a local directory for the session? All right, here we are running Open Code connected to the remote server. We don't We don't want this. Uh but in the meantime, let's open up a second terminal here. And do Let's same thing. Come on, place. Same thing except a different project directory. Okay, two connections here. Um Do this. I'm just doing some example queries.

And um let's have a look here. If we Um no, we have not caused enough load to scale that scale and increase yet. Um Let's see. >> I don't think um I should have set up an automatic load. Uh I don't think this will actually trigger a scale up event. But at any rate, uh you saw how we were easily able to define a inference service that would scale with um with load and connect um open code instances running locally on your computer to that um auto scaling inference service. >> Okay, so just to recap, um we showed you we we talked today about AI, uh you know, using FuzzBall and with a particular focus on using FuzzBall to uh do sovereign AI workloads.

And we showed you two demos, um they were really pretty cool. Uh the first of which was FuzzBall running on a DGX and um actually, you know, we're dog fooding FuzzBall using it right here within CIQ to solve real problem that we've got. And that is to produce, you know, more training videos, to produce more promotional videos, and and produce more content. And we had a really fun conversation with a Data Hara talking about um how he joined the team and was kind of tossed in the deep end and was able to swim very quickly because FuzzBall is so easy to use. So, you know, appreciate him coming on and and talking about that.

And then we talked to Wolfgang Rasch and he showed how you can take these inference models and you can use FuzzBall to stand them up really quickly and to be able to scale them up and scale them down as you need to in order to allow a large number of people to use these models. All right, cool. Um so, this is just like, you know, FuzzBall is a really multi-purpose tool, right? It is a It is the Swiss Army knife of HPC and AI. And there are all kinds of things that you can do with it. So, this is just kind of like a little fraction, a little bit of what you can do with Fuzzball.

And we're going to be, you know, showing more demos in the future and doing more webinars and talking more about, um, the different things that Fuzzball can do. If you have topics that you're interested in, um, if you want to know more about Fuzzball, you can go ahead and email us at info@ciq.com. Or if you have questions that you didn't think of, um, uh, you know, during the recording here that you want to you want to, uh, talk to us about later, then that would be a fine thing to do. You can You can email us there as well. All right. So, thank you for attending.

Um, hope you had a fun time. I'm sure I did. And, uh, we'll we'll talk to you later.

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert