
Sovereign AI and interactive HPC: unifying training inference and exploration in one workflow
Most AI teams run training on one platform and inference on another. The handoff between them happens outside any workflow definition—it's manual, fragile, and repeats with every model iteration. For organizations that can't send proprietary data to external AI services, this fragmentation isn't just inconvenient. It's a blocker.
In this webinar, we demonstrate Fuzzball service endpoints—a capability that unifies AI training and inference in single, portable workflows. This is sovereign AI without the complexity of managing separate platforms.
- Jonathon Anderson, Principal HPC Product Engineer, CIQ
- David Godlove, Product Engineering
- Eric Hendricks, Technical Product Marketing Manager
RESOURCES:
Fuzzball product page: Fuzzball Fuzzball Documentation: Fuzzball documentation Request a demo: Request a demo
Press Release
Blogs
- AI workflow orchestration: why separate platforms fail
- How to run interactive HPC workloads alongside batch jobs in a single workflow
- Running Jupyter notebooks on HPC clusters without SSH tunneling
Transcript
Good morning, good afternoon, good evening, wherever you may be hailing from. My name is Eric Hendricks and this is CIQ's first webinar of 2026, if you can believe that. I want to welcome you all today. Today we're going to be talking about an exciting topic built around fuzzball. Uh we're going to be talking about sovereign AI and interactive HPC unifying training inference and exploration in one workflow. So that is a lot of words and I could tell you all about it but it'd be far more interesting if I brought a few people in to tell you about it for me. We'll get a little bit of a conversation going on.
Uh but before I introduce my guest today, a little bit of housekeeping. Uh we are live. If you haven't noticed over on one side of your screen, it's either that side or that side. You uh there is a chat box. feel free to uh to say hello. Let us know where you're uh where you're watching from. And if you have any questions during the course of this event, feel free to put your questions in there. Um and I will make sure that uh one of our amazing guests will um will be available to answer those questions. We'll answer those live on air to the best of our ability.
Um I really I think that's it. So if uh if you haven't, make sure to visit our website, sign up for our newsletter. that way you get notified anytime that we do these events. We are planning on quite a few of them uh covering uh fuzball, Ascender Pro, uh Rocky Linux, all the things uh that CIQ is is uh passionate about. So really looking forward to this. And with that, I will shut up for a moment and introduce our guest. Dave, would you like to go first? >> Yeah, sure. Um thanks very thanks very much. So, uh, my name is Dave Godlo and, um, so I'll just tell you a little bit about myself.
Uh, and you know, so my background is that I used to be a neuroscientist working at the National Institutes of Health and I became interested in high performance computing while I was there. And so I joined the Bowolf team as a staff scientist um, working at the intramural uh, HPC resource at the NIH. That was about 10 years ago. It was around the time that containers really took off within HPC. And so I've been kind of an HPC container nerd ever since. And uh I work with the fuzzball team primarily and I am a um I am a technical product writer but I also do a lot of HPC engineer um testing all kinds of stuff.
View full transcriptHide full transcript
>> Awesome. Well glad to have you today Jonathan. >> Yeah so my my background is in HPC as well but more on the systems side. So I don't have any uh any research science or academic background there. Uh but my job historically has been to to support research and scientists using an HPC system and I've done everything from the queueing system to the network to the storage to the Unix systems that undergur all of that and uh joined CIQ like three and a half four years ago something like that now it's getting a little bit um and now my my role here at CIQ is uh as a product engineer so I'm primarily responsible for our HPC products which encompasses fuzzball but also werewolf and werewolf pro and to a slightly lesser extent, but still part of the conversation, Appainer, um, largely as part of that werewolf, uh, product stack.
Um, and yeah, really excited to talk today about new features that we've been able to add into Fuzball that have I think have been a long time coming. They're they're filling a pretty big need in the product. >> Awesome. Well, thank you both for those introductions. And I realize I don't know that I introduced myself. This is uh my my first appearance here on the CIQ channel. So, uh, real quickly, my name is Eric Hendricks, uh, known online as the IT guy, and I've been a Linux systems administrator. And then around 2019, I dove into the dark side of sales and marketing. Found my love of teaching and and content creation.
U, and so now I spend a lot of time talking to other other CIS admins and nerds about what it means to work with Linux and containers and other similar technologies. Uh, so if if you've seen my face, it's probably because you're used to seeing me at this angle, um, or conferences. So, uh, really excited to be here. I've been part of the CIQ team for a whopping three weeks now. Uh, and they felt comfortable handing me a microphone and and a webinar. So, we'll see how this goes. >> Welcome. [laughter] >> Yeah, great to see you, Eric. >> Yeah, good to be here. Good to be here.
Um, so, Jonathan, I want to throw the the first topic your way. U, so right now, uh, folks that are getting used to this idea of hosting their own AI, so we're not talking about chat GPTs here or clouds here. we're talking about. We need something a little bit more specialized. We need something maybe that's sovereign, something that we host in our own data center, whe that's cloud or hybrid or local. Um, but right now there's there's kind of two aspects of this problem there. There's the training aspect, training the model to know the things we needed to know and the inference side. Um, and right now the industry sees those as two very very separate tasks.
So kind of talk us through what kind of let's let's flesh out that problem. Let's define what what we're trying to solve here uh between sovereign AI training and inference. >> Yeah. So a a big part of CIQ's uh positioning for Fuzball and frankly most of our products maybe all of our products is this idea of you the customer having complete control of your infrastructure and your data and everything that you're doing. And that's a little bit at odds with the way a a fair bit of the industry is going right now of hey give us your data, give us your your systems and run it on our things and uh trust us it'll be okay.
So there there's a time and a place for that. Um and you know I'm using cloud all the time right now. Uh but depending on what data you're wanting to process with an AI model, you you might either be prevented from doing that regulatory or you might just not want to. you you might have uh you know business secrets or whatever the data that you don't want to leave your premises or your control. So our goal is to enable that full stack um to run in your infrastructure with your systems and your infrastructure may still be the cloud. It may be some contractual system that that you uh run outside in AWS or something but it may literally be uh equipment on your premises.
So, where this comes up in the AI conversation, uh, you mentioned training. I want to break that down a little bit because that that gets to be a little bit of a fuzzy term. There's training of the the model itself that you do on large corpuses of data uh for weeks or months or longer. Um, and that is certainly a thing that you can do as a traditional batch process and could conceivably do on fuzzball. It's not necessarily what we're targeting with uh with what we're going to show off today. It's certainly not the part that's enabled by the new functionality service endpoints that we're happy to show off.
Um, Fuzball has been from the beginning a batch processing engine that can do large parallel work and you could do that um to or you could use that capability to train a model let's say uh of various sizes. But there are other parts of um what can be thought of what people might think is training that are are not quite that same thing. You might fine-tune a model. You might take an existing pre-trained model and then give it additional data about your specific data set, your specific concern, your specific industry um and and retrain that model with uh more information to be specialized for your use case.
That's one thing that you can do that you might want to do. The other is um uh through various means bringing data in into a a vector database or a rag database that will give that model access to information that it might want to look up during its execution. So when it's inferencing, when it's answering a question, let's say when you're asking it to produce output, um a database like this gives it access to well maybe I can keyword search and and get fuzzy information about uh about fuzzball. That what we tend to do is is show off asking it questions about Fuzball. And I'm sure that's the thing that we'll do here.
um ask it questions about your product or your support database or whatever information you've brought in and and it can answer questions about it like you would if you were accessing that information. That's a a a computationally intensive process as well because there is an inference step in looking at the data and deciding what keyword searches let's say for example that's not quite how it works but you can imagine it what searches the model might produce into that database and what results it should get out building that index is is computationally intensive as well that's also a bit of a batch process but then the final step is that inferencing step which you can do as a batch if you have a whole bunch of prompts or you're going to generate a whole bunch of images or something like that.
You could do that all as batch work as well. But more commonly, people want to do that as an interactive process. You might do that initial ingest, that initial retraining as a batch process, but they want it to terminate in a final model that's running as a service that you can ask questions of, you can log into and uh and and just use in your day-to-day uh work. So that's that's the new capability that we've introduced into Fuzball that you can put both of these kinds of work together in one platform that a research scientist or an IT person like me or really anyone could run with the platform and and use.
>> That's awesome. Yeah. And I forgot to mention that the reason we're why we're doing this topic today is we just announced service endpoints. Uh so that's uh that's kind of a big deal. U so appreciate the explanation, Jonathan. that that really uh that really helps kind of clear things up a bit. Dave, I'm kind of curious uh with with your hands-on practical background, is is this something that you wish you would have had? You know, what what are your thoughts? >> Yeah. Yeah. So, to get back on the on, you know, talking about service endpoints and that being kind of the impetus for this work.
So, um it it's kind of funny because a lot of times when we introduce fuzzball, what we talk about is we say, well, it's it's um container orchestration, right? And that's really what it is at its heart. And so then a lot of the next logical question is like well how is it different from Kubernetes and the answer to that is well it's an entirely different mindset because Kubernetes mindset is focused really on services and taking multiple different microservices and sticking them them together and allowing them to talk to each other and do things like that. Whereas Fuzball is focused on jobs.
And so with that explanation then it becomes kind of surprising when you go back and say and now we're doing services but the the the the reason for that is just what you said you know like so it is we we've found both through experience and also through talking to other HPC engineers and so on that um you know users of HPC do need services right I mean if you think about um how many of them want to run Jupyter notebooks for instance or want virtual desktop environments or want to visualize something with parview or something like that. But these services, they're still a little bit different than the kinds of services that you would normally be thinking about when you were thinking about Kubernetes.
um they're not necessarily services that need to sit behind uh a big kind of networking stack with ingress and egress and all that kind of stuff that they don't need to necessarily have like load balancing and um high availability. Basically, a user needs to spin up a lot of times a server and then connect to it and or have a job connect to it and interact with it. And that's essentially what they need to do. And so that's kind of the way that we're thinking about services within Fuzball is once again from an HPC users perspective. And this, you know, once we started thinking along these lines, it really, you know, with with the with the sovereign AI idea, it really kind of began to blow up and say, "Oh, wait a minute.
There's all these things that we can do as as just as Jonathan was saying as far as like training and augmenting existing models and then serving them out all within the same workflow." Um, so I think that's that's all those topics are really what we're going to be kind of covering and talking about today. >> Love it. So Jonathan, back back to you. Um, so we kind of talked about what the problem is. So what where where does that bring in the solution? How do we fix this problem? >> Yeah. So for for fuzzball. So I I kind of skipped over it's because we're talking about fuzball all the time.
It's sometimes easy to forget to just introduce what is fuzzball. Fuzzball is um CIQ's vision of the the future of high performance computing in a massively uh um hybrid environment. So you might have some resources in the cloud, you might have some resources locally. You want a single interface that lets you operate all of your compute wherever it is and you want your workflows to run the same no matter where they run. So there are a number of technologies that go into this. We have a a workflow definition that talks about how one job flows into the next and how their inputs and outputs interact.
We have uh all of our jobs are containerized so that they bring their dependencies with them and so that if I run my job here on this environment, it's the it has the same code and same libraries as if it runs in that environment over there. But what it's been historically is it's batch processing. And you can you know we we have a a web interface. It's all nice and easy to use. Um that can show you a block diagram of your environment, but uh or sorry, of your workflow. Um but it's been exactly that, a batch processing system where you have a batch job that runs in a batch job that goes to another batch job.
And you can run multiple jobs in parallel and then a a final job might depend on multiple outputs from multiple previous jobs and it can break it up and parallelize it that way. But for this final result where you want to monitor um you want to monitor your job in real time, maybe visualize it with something like Parave or you want to run a desktop suite, you want a virtual desktop. Um or our kind of prime case here, you want uh an AI model that's going to run and and be interactive over time. uh that is something that we've had to in the past kind of simulate where you you ran a a normal job and then you had to port forward into it either using SSH or using the fuzell command line and service endpoints alleviates that.
Now in addition to having batch jobs in your workflow, you can have a service in your workflow and that service can have the same relationship to other jobs that the jobs themselves have with each other. it can depend on other jobs having completed or being uh running right now and then you can log into that service using an endpoint. Um technically we can have multiple endpoints in a service but I don't think we've done that. Um, and so you can click a button in the web interface and get access to a virtual desktop that lets you interact with something that's running in parallel. Or you can push a different button and get access to a Jupyter notebook and do some data science there and manipulate a data set that you have in real time.
Or uh you can press a button and access an AI system and an AI model that's running and and ask it questions just like you would with Claude or ChatgPT or something like that. And don't underestimate the ability to add a lot of this configuration into into a tool like Ascender and using Ansible playbooks in the background. Uh you can further automate a lot of this uh using ascender pro. >> Yeah. So we we tend to use ascender as a a deployment mechanism for fuzzball depending on the environment for sure. Um we don't have anything in the front yet that runs things in fuzzball with ascender.
We should do that. That could be cool. Oh, yeah. I'm just reading this post here from Gwen. Thanks, Gwen, for putting that up there. Let's see. Okay, so we we kind of talked about the this problem between uh training and inference. Uh we we've kind of talked about how what what fuzzball is and how uh how fuzzball can kind of help this solution. Um, but some of our audience may be sitting out there wondering uh is is this a problem that I'm having? Is is this something that I can uh really contribute to or is is this is this a solution that I can uh words are hard today.
Is this a problem that I can is this a problem that I have? Um so let let's talk about some specific use cases. uh when when would uh when would sovereign AI uh and some of these tools like service endpoints really come into play? >> Well, it might be easiest to to make that case just by pulling it up and and walking through some use cases and showing what that looks like. Are we ready to do that or or um what do you think? >> Uh Dave, you have anything to to contribute before uh I forgot to absolutely get to tell everybody we got uh we got shiny demos today.
>> Yeah, we do. Um yeah I I surprisingly I do have something to contribute today. Um [laughter] so uh yeah yeah yeah no I I think that audience members might be having that that kind of question. They could be you know thinking to themselves is this a problem that I have that is going to be you know solved by um fuzzball that I can I can use. Um, and I think that a lot of the times, and you know, we know this, but it's good to be reminded. A lot of the times when you've got new technology that you're you have access to, um, it doesn't it doesn't only solve problems maybe that you were thinking about that you already were aware of.
Sometimes it solves problems that you you didn't you didn't even know that you had. And sometimes it doesn't really solve problems, but it opens up new avenues for research that you didn't realize that are just like, oh, I never even thought about trying to do this, and now I can do this. I can incorporate visualization for instance into the middle of my um my HPC workflow and I can kind of pause it and go in and select some stuff and then resume it you know based on and I can you know analyze data based on the regions I've selected in this particular model or something like that.
So yeah it it really um I think once uh we start kind of showing it off here uh momentarily I think it's really going to uh open a lot of eyes and it's probably going to spark a lot of ideas. I will say that the the addition of this functionality has been the first time that I've been this excited not only to be you know we use the term dog fooding using fuzzball within the team within CIQ itself even though you know we aren't HPC people but suddenly it becomes useful for a much broader uh set of things that we keep making the case computing is high performance computing now uh right most of what we're using our computers for certainly all this AI stuff is performance sensitive.
You have you need real horsepower behind it. So, I'm I've been excited to use it uh kind of as part of the development process, but I and and other developers on the fuzzball development team have all expressed like, oh, I I have I could use this at home. I could be doing this cool internal sovereign AI stuff, you know, just as a hobby thing, too. And Fuzball is enabling that. But, you know, part of the point of it is that it's scalable to all of these use cases. We wanted to scale down to to local development and and interesting small things, but then scale up to to large production scale work also.
And and this is definitely a step towards that. >> So what you're saying is I should use this technology to run my sovereign AI that serves as my codem for my Dungeons and Dragons group. Is that what you're saying? >> You you probably should. Yes, >> you should. You totally should. >> But but I can also see like a a future, you know, I mentioned I'm I'm using Claude. use uh Claude and other AI tools internally uh as part of development. But there's a future where maybe you don't want to have your code exposed to an external AI model. And what if you could run an open uh and community accessible model or even a proprietary model that you buy and then run internally on your local infrastructure and don't have to expose your code out to a third party.
Um that that's the future we're trying to enable here. It's of course computationally intensive but even beyond just having the equipment for it uh running the infrastructure can be uh its whole own thing and uh we're hoping to make that easier. >> Awesome. Well, I I'm sold. You you can take my money or we can just jump into uh into the demo. What do you think? >> Uh we can do both. I'll I'll take your money and we can give a demo. [laughter] >> All right. I got your screen up here, Jonathan. What are we looking at? >> Great. So, I just realized my computer is not plugged in.
That could become a problem later. Um [laughter] um so this here is the the fuzzball web interface and we're right now looking at the workflow catalog. This is one of the the things that we're really proud of with Fuzball that because the workflows are portable and run the same regardless of where they run. uh we can build an ex a a a workflow for you and you can run it in your environment and it will work the way it worked when we developed it. And you know these these are templated. So when you run one of them uh you're asked a series of questions if there are questions for it um on you know what inputs and outputs you want for it maybe what storage you want to provide for it depending on the the workflow.
Each can have its own uh its own inputs and outputs. But this this alpha fold one for example, I think. Yeah, that's this one that I've got pulled up here. This is that block diagram that I mentioned. And we can zoom out. Oh, not even far enough. This one's really wide. Um, but this is a series of batch jobs, right? And most of these are not dependent on each other. So they're running kind of whenever they can, whenever there are resources available. But then there's a final job that needs output from all of them that runs a final prediction apparently. And uh then you know this is all batch work and it concludes and your workflow is done and you get an output.
Uh but what if you wanted to be able to view the output of that either when it's being uh generated or kind of monitor the process of this. You can of course monitor the workflow itself through the the fuzzball interface, but you need some kind of bespoke interface. Um that's where the service endpoints come in. And so like the the most obvious example of this, something that you want to run and be uh interactive with as you go is a virtual desktop service. The the icon for this one is a little bit more generic right now because I've had to modify it to support service endpoints.
We we just got that working. Uh but we have, you know, this workflow running and I have some data in a volume that's persistent between runs. I can mount whatever data I want to it. It's running this container that we have that runs a a virtual desktop and uh using XFCE. It could be whatever you want, but this is a lightweight one. And because this is a service, you see it says persistent service. It's going to keep running even when there aren't other jobs in this workflow. Something we haven't actually done much of right now, uh is an inworkflow service. It's really only meant to be running when other jobs are running because maybe the jobs need to use it.
Um that's a thing you can do with this. That would be an ephemeral service. But this persistent service is something that will run forever until I stop it. Uh unless there's, you know, you can have a time limit on it or something like that. But I've got an endpoint configured here called desktop that I can just press this button on with this running and connect to this VNC instance. And here I it's running from before because I got this all set up ahead of time. But I have, you know, some some graphical stuff running. uh high frame rate in my GLX gears of course and uh I can browse the file system.
I can see where that data was mounted. I think the data is just mounted to the home directory here. So that that would be persistent. Um, but then you know you you would use this for example if you're using a desktop suite like the Ansis suite of software has a a workbench uh product that needs to be run in a desktop. And so you could use this to run that on a shared infrastructure. By shared I mean like multi-user infrastructure that you control, but it doesn't have to be the desktop sitting in front of me. And then that can interact with other parts of the fuzzball cluster as well.
Another popular uh system for this is Jupiter. The Jupiter project grew out of the Python and IPython communities to do um interactive notebook computation where you could have some documentation and markdown and interle that with um uh interle that with co code blocks that could execute and generate output in the notebook and then you could package that up and and save it. Running these, however, on an HPC environment is kind of tends to be that thing I I described earlier where you you run a job and then you SSH port forward to it and and access the thing that's running or you run it locally on your local workstation.
But with service endpoints here, I have a persistent service for Jupiter just like that virtual desktop. I can click this connect button down here and all in my browser uh I can access the notebook that I had before. or I I actually it's it's actually kind of cool to just see how easy this is to do. So, I'm going to delete this uh right here and just drag this notebook back in. Oh, if I can find it. Oh, no. I think this is it. Yeah, right here. and just interact with it as though it were running in any other Jupiter environment, but have it be mounted to my external storage that is part of my infrastructure and dispatch from here onto multiple nodes that might be running in the background as part of the same workflow.
Um, and then execute the command blocks here. So I can say uh restart kernel and run all cells. And we can see this Python code, you know, over time executing the code that's in this notebook. It takes a little bit to initialize this first time where it's bringing in mapplot lib. Um but yeah, here we go. And it'll generate some mapplot lib graphs and and things like that. So that's uh two of the the really basic examples. Um the next one then is the AI workflow which I would throw to Dave for uh unless we have more to discuss before then. Well, I maybe before um jumping over to me uh you know it might be worth kind of mentioning and tie back to that workflow catalog.
>> Yeah. >> That like Yeah. I mean this this Jupiter example that you just showed I mean it's just me. Um [laughter] >> this Jupiter example that you just showed um it it's really cool because it's in the workflow catalog and you can just go and click on it. Yeah. >> And run it. you don't have to, you know, come up with any kind of a workflow. Um, you know, we call them fuzz files. You know, the ammo files already saved within Fuzball. And that's something that with your new installation of Fuzball, it comes already prepackaged and ready for you to go. So, your users can already just open up the workflow catalog and clickity click and be in a Jupyter notebook um environment um coding things up.
>> Yeah. And we're adding entries to that catalog all the time. So uh the the Jupiter server example that I was running here completely unmodified. One of my co my colleague uh Wolf Gang um who who's kind of waiting in the wings here developed that and I didn't have to do anything to it at all. I loaded it up out of the catalog, clicked run and it ran for me just as well as it did for him. The desktop one I had to do some development on because it hadn't been ported to services yet. In the past we had provided that in the catalog but you had to do the port forwarding work uh to actually access it.
and when it ran, it would produce some log output that told you how to to get that port forwarding working. But with just a a small little addition to the YAML, now Fuzball understands that that's a service that's running. And so you can connect to it with the web interface and we'll get that back into the catalog so that all of our Fuzball customers can access that and and do exactly what I just did. Have a a virtual desktop running wherever their fuzzball instance is running without uh having to to do any work other than click run and click connect. >> That's really really powerful.
I mean you think about having disperate teams remotely you can you can spin up VDI desktops for your entire team whether they're the scientists that are in the lab or whether they're maybe traveling scientists that uh that visiting universities or other research laboratories they can access the same hardware uh no matter where they are and take advantage of that that powerful compute system that's running underneath without having to send multi,000 laptop tops out into the field. >> Yeah, there there is one thing I want to call out here because it's it's easy for us to miss it again because we're using the system all the time that the the capability to develop something like that is as accessible to every member of your team also.
So there's no gating on, you know, anyone could have put that Jupiter example together because they just bring a container and they write the workflow and it runs and then they can share it with anyone else on their team or anyone else in their organization. Uh and because it's portable, it runs the same for any of them. Some systems like this, there's a kind of a hard line between publishing applications that are available for other people to run and the people that can run them. and and you kind of can't get an application into the system without administrative support, without giving that work and expecting that work of someone else.
And Fuzball provides all of the power of that uh all the capabilities of the system to everyone that's using it, but not everyone has to. you can pick something off the shelf and use it. But if you're the kind of person, if you're a power user who or you're a grad student that's been tasked with it, you can uh you can do the work also and then make it available to others, which it's really cool. >> That is cool. Uh so Dave, I think you had a few things you wanted to show off as well, right? >> Yeah, absolutely. Yeah. and and yeah, if you could um share my screen and and and just to further emphasize that point that fuzzball is really facilitating collaboration.
So, I'm going to give a shout out once again to my colleague Wolf Gang who um uh Jonathan already mentioned before. Yeah, go ahead and and share the screen because um the workflow that we're going to be looking at here, this is something that actually Wolf Gang put together and he actually started for me uh yesterday on and and I can just go in and look at this um this workflow myself just by virtue of the fact that we're both part of the same engineering group. And so I can go in and I can have a look at this workflow. I can see the fuzz file that he used to create the workflow.
I can open it up in the workflow editor. But I'm going to go through and I'm going to talk about this workflow in a little bit more detail because um now we're going to be getting into talking about uh sovereign AI what we what we kind of like you know started off talking about what we you know kind of teased at the beginning. So I think you know the the workflow dashboard is is good once you kind of know what your workflow is doing and you're like monitoring each of the pieces of it. But when you don't really know yet I think that the workflow editor makes it a little bit easier to kind of see your workflows and see what they're doing.
So, I'll pop over here and we'll talk about it at a really high level and I'll give you kind of a sense of what's happening here. All right. So, backing up, when we first started talking about um using Fuzball for AI, one of the things that we started to do is to is to say, well, what what would companies or organizations or labs or whatever want to do with AI within fuzzball that they couldn't do just, you know, without fuzzball? And one of the things that we started to think about right off the bat is like well what do we want to do?
What do we want to do with AI in our company and you know I think that that's a really good starting point just think about what we want to do and try to solve those problems and then that's going to you know dog fooding it like that is going to um you know make the make the uh the use cases kind of pop out.
And so one of the first things that came to mind obviously is let's build [clears throat] a model which has information about how fuzball works and you can talk to the model about fuzzball and you can ask questions about like you know what do I need to do in order to do x y and z with fuzzball right and so how would we you know and that's a great use case too because we're developing fuzball and you know sometimes there's new features in fuzball sometimes these are things that we don't want necessarily to have out in the open yet. They haven't been released yet. So, we want to be able to host all this in a sovereign AI kind of way.
We want this to be kind of like internal to our organization, but within our organization, we also want to be able to like share this out and have new employees come in and be able to ask questions about like fuzzball and get up to speed really quickly and things like that. So, that's exactly what we're doing here. Um, so we've got this collection of services here in blue and jobs here in green. And I'll kind of walk you through and and talk a little bit about what these services are doing and what the jobs are doing, what the dependencies are between them and so on. So we're using this thing called local AI.
And local AI is an open-source way to host uh your own your own models. And so within here, we're hosting a couple of different things. Um, we're hosting um kind of like the open source um the open source version of chat GPT4. And we are also uh hosting um a a service that allows us to do text embeddings. And I'll get back to that and talk about that in a few minutes. And then we have a couple more services which depend on this one. So you can see with the dependencies set up that this one has to start first and then these ones can start. We have one here called local recall.
And what local recall does is it allows us to serve a vector database. So you you've probably heard about people talking about rags, that is um um uh recall augmented generation. Um so so essentially we can take um text and we can vectorize it. We can put it in a database and we can use that to improve our queries when we talk to a large language model. And so that's what this is going to allow us to do. And then we also have here this uh local AGI and this is like for um kind of creating agents and making your models agentic and I'll talk a little bit about more about that.
I think a lot of people kind of already understand what agents are and what agentic means but I'll I'll talk about that a little bit when we start talking about this. And so here is where you know up until now if you've been following fuzzball progress you've seen us you know show these these workflow diagrams and they've been these directed as cyclic graphs and you can kind of think of information or [clears throat] data or whatever flowing through the graph from the start to the end. Well, you know, these this still is a a directed a cyclic graph, but I think it you have to change your mindset a little bit because now um what we have is a couple of jobs here at the end.
This one is just sort of like a little checkpoint job to to say, okay, do I have all my services up? Is everything healthy? And then if that is the case, then we can proceed on to this ingest corpus job. And what that's doing is it's going to our private GitHub repo where we host uh the fuzball, you know, fuzzball code as we're developing it and it's grabbing the markdown files that make up all the documentation and it's shoving that back up into this local recall service which is up and running and allowing it to create that vector database that it can use for the rag um sitting in front of the local AI here.
So that's that's how all this is working. So, you got to kind of change your mindset a little bit. It's a little different um because we're not, you know, information is not flowing. And just as an aside, um what the actual uh the actual data that we're fuzzing that we're putting into this model is actually hosted right here within FuzBall as well. So, it's it's this documentation and um I'll just give a little plug too while I'm here. you should know that this documentation is open. Um maybe somebody can flash up on the screen or put into chat um the place where you can go to get this.
Actually um this is a development version of the documentation. So I won't put this link in but maybe later on we can share the link out for the uh the version which is um not development. Jonathan just did. So all right cool. So let's play with this a little bit. Let's go back to the the actual running. So before I do that, is there have did I forget to explain anything? Do you guys want to jump in and clarify anything or >> ask any questions? >> Um maybe the only thing I'll talk about so that that uh that local recall is is the bit that I was talking about before with um you know it's it's it's maybe a little bit surprising how computationally intensive that is if people aren't familiar with how it works.
And that's where you're you're using an AI model to kind of get the the meaning of content that you're putting in it and index it by the meaning by the semantic meaning rather than just by keyword search. So that's its own AI workflow all by itself which is kind of neat. It's been good. >> Yeah. >> And just think about this too, how powerful this is from the perspective of any company that has a bunch of internal documentation that they want to their employees to be able to go through and look at um or talk to an AI model about.
And the other great thing about this, since you're controlling all this yourself, as your documentation changes, you know, you can trigger jobs to go ahead and update this um nightly or weekly or however often your documentation is changing so that it's always up to date with the latest, you know, whatever. [snorts] So, let's go back in and play with this. So now we've got um it's kind of a little confusing because we've got these these jobs here at the end that have have finished, but we still have services that are running and we're we're in a situation now where the user me can interact with these services, but we're also having jobs that are interacting with these services and the services are also interacting with each other.
So I'll just kind of jump in and have a look and kind of show you what's going on. So I told you this local AI is a way to host your own uh models. They could be large language models or you know uh image generation all kinds of different things but this one is a is a language model. So it's uh GPT4. So let me jump into this model and say okay um tell me about fuzball and its rel oh it's relation to HPC. And if I do that, it's going to have no idea what I'm talking about, right? It's going to say, "Well, fuzzball string theory," which is cool because that's actually where the where the um name comes from.
But that's because this is just the model itself with running without any rag in front of it. So this is what you would get if you just went up to chat GBD, you tried to ask him about fuzzball, right? So let's go back over here and say okay. [clears throat] We have this local recall service here and that I can connect to as well. Um it's not really useful for me to do that because I've already you know there's already been a job which has uploaded the data to this and it's already created the vector database but I could do something you know manually interactively if I wanted to.
>> Can you see it in here? Can you see that it's been put in? I haven't interacted with local recall directly yet. I Yeah, here it is. >> Yeah, nice. >> And yeah, I don't want to mess with it because I don't want to mess anything up. Um, but here's where things really get fun. So, this local AGI, so this this once again is um this is a a service to try to make your AI more agentic. And so, what do I mean by that?
I mean what I have always heard or what I heard several months ago uh when people were talking about this a lot more is that the idea is you have an AI agent and that agent is like your coworker right you send the agent Slack messages and they do stuff on your behalf and maybe they save it in Google Docs or they create GitHub repo polls you know PRs to your GitHub repo or they um send you emails or something like that or maybe even meet them in Zoom like this and talk to them than Zoom. So this this is a a service to allow you to take these models and make them more agentic.
And so we've we've kind of done so Wolf Gang has kind of done a little bit of that already. So we've got this fuzzball agent here and I can chat with this fuzzball agent and I can ask it questions like, you know, tell me about fuzzball and its relation to HPC. And if I do that, you know, it's going to take a few minutes as it grinds away on that, but it's actually going to go and search the vector database, and it's going to use that to augment this this um query that I've given it and give it something that's a little bit more actionable that it can talk to chat GPT about.
And so, um it's it's taking a little while for this to work. This is the longest I've ever seen it take. So, hopefully, you know, the curse of a live demo. >> Yeah, I was really happy. There we go. showing it off earlier doing. >> Wow, it gave me a lot. >> So yeah, so it says Fuzball's high performance computing 2.0 enables you to create and run workflows in HPC. So it knows about workflows, various stages such as storage volumes, fetching containers, all that kind of stuff. So this thing actually knows about fuzball now because it's got this rag is sitting in front of it. [snorts] Um, so yeah, so I could ask it more things about Fuzball, but in the interest of time, let me go back and show you some of the cool things you can do with this as well.
So, um, [clears throat] if I were to create an agent. So, basically, this allows me to really easily do things like connect an agent into something like Slack or something like a GitHub PR or something like that. And so, let's say I connected it to Slack and GitHub. Then I could go over here and I could set up filters and I could filter on specific like reaxes for instance. I could say hey within GitHub PRs I only want the agent to be able to listen when I say at fuzball agent or something. I could create a reax to you know to point to that. And so what I could end up with is an agent that I can talk to in Slack and I can ask to go review PRs or go look at code or you know do things like that.
Uh you know based on this fuzzball based on this fuzzball workflow that's running here. >> And to just hammer it home like it's it's all running on our infrastructure ours because it's our fuzzball. If you're running fuzzball it's running on your infrastructure and your data is never leaving your premises. It's never leaving your environment. Your your data is in that database that's running that's backed by storage that you control and you own the entire stack. Absolutely. And let me show you this as well. So, um, another thing that we haven't really talked about, if I can find it in here, it might be easier to find in a smaller, uh, smaller one, but um, so there there's a lot going on here behind the scenes.
Um but if I go here, I can see like so so these these these are served out using um so so the the services themselves have some new um some new sections within the fuzz file that allow you to configure like what port they should be listening on, what kind of traffic should come through. They have this scope here. And that's one of the cool things about these is you can set the scope either to just yourself and say, I want to be the only one who accesses this model because I don't want Jonathan to know about how I'm talking to my model about Dungeons and Dragons, for instance, and you know, I don't want them to know my DM prompts.
>> You don't want your players looking at it. >> Yeah. Exactly. Exactly. Or you can have um you know, other people in your group. So, we've got this engineering group here. Or we can have like at this at this case the entire organization. So anybody who's got access to this fuzzball organization can access um these uh services and they can go in and start chatting with the fuzzball agent about fuzzball. >> Yeah. There there's also um a public scope for services that you actually want to serve with fuzzball as a production service outside of you maybe outside of your organization but in any case outside of fuzzball itself.
So we you know this is new stuff. Um, but we the goal is to support kind of that entire stack of internal services that you're using for your de development and research right now all the way to production services that you want to make available to uh maybe the rest of your company or beyond. >> I didn't even know that we had that kind of scope. That's that's that's really cool. >> Yeah. >> Like like like you said, this is really new. It's It's been GA for what, two, three weeks. >> Yeah. For various definitions of GA. >> There there you go. Fair enough. Um, so I wanted to uh we're we're getting close to time here, so I wanted to to kind of shift uh a little bit if you don't mind.
Um, we've we've got three different questions that I'd like to uh to address from the audience. >> Awesome. Um, so Sean asked, "Is it possible to restrict an application or iteration of an app within Fuzball?" >> Restricted to what? I guess >> I presume that's the text we have. I I would say that like the the application like you are in complete control of the application. So at any time when you start it, the workflow doesn't have an access to any data by default. it only has access to data in volumes that you mount to it or um or objects that you ingest as part of the workflow.
So you you have complete control of what it's able to see because it only sees what you explicitly give it access to. >> Yeah. And and yeah, if you're talking about restricting it to who can see it, once again, you can restrict the scope to yourself or to groups or to the organization. Um, also you know we you're able to ask for specific resources. So you can restrict the CPU usage, the memory usage, the type of GPU that it gets, the type of networking that you set up. Yeah, you have very fine grain control over everything that you can do with it. And you can also give it timeouts.
So even if the thing goes crazy and you know keeps running um you can say don't run this past 2 hours and fuzzball will just just like you know any other batch scheduling uh system it'll just cut cut it the uh cut it short at however you know whatever you set the wall time at. >> Yeah. And you can build all that up as either a textbased YAML document that says what all is in your workflow and how it relates to each other, how the services are configured, how the data gets ingested or or mounted into it. Um, or all of that you can, you know, clicky button do in the web interface or both.
You can put it in YAML, bring it into the web interface, modify it, and then save it out to YAML again later if you'd like. So, awesome. Thank you guys for for that. Uh, Sam asks if uh any part of uh the AI demo that we were showing off today is using uh MCP as well. So the the various components of the local AI stack are using MCP to talk to each other. So the local recall service is an MCP server and the local AI service that runs the chat model is an MCP server and then the local AGI application is an MCP client that sends prompts to those various services.
So in effect it it tells the uh it you you type your chat query and it amends that query with answer the you know it gives it the prompt but then also tells it what other MCP services it has access to and so it generates additional uh MCP calls that the local AGI system puts to local recall or other MCPs like the Slack stuff and the the GitHub stuff that Dave was showing off there. Yeah, I maybe should have said at the start too that um this is the the the actual system that we're using here is is local AI and that's built to be a drop-in replacement for Anthropics Open AI.
So um or yeah for open AI. So it so basic basically the API um the API is is supposed to be exactly the same. >> Yeah. >> Awesome. Great to know. Uh, Alex asks, "What versions of Fuzball is this enabled?" >> Uh, this is available in Fuzball 3.1. Uh, and you know the the subrees of that and beyond. Let's see a couple more comments coming in. Let me double check and make sure. Uh so Hillilmar I I apologize if I butchered your name. Uh how do workflows fit into classical concept of resource resource management on shared HPC resources? Is there u stuff like a scheduler under the hood that would take care of managing multiple users competing for compute resources?
>> Yeah. So fuzzball is meant to compete in that space for sure. So it it has auler. Thatuler is a little bit different than a traditional HPCuler because it also has a a resource provisioner. So it natively understands the idea of oh I don't have resources available but I have access to cloud resources over there. I will start an instance and then run jobs over there and then dep prior or deprovision them when they're no longer necessary. But then it can also uh schedule resources on or schedule jobs to run on um local kind of on-prem resources. All of this works using a uh container runtime called fuzball substrate which I like to describe as kind of slurmd plus aptainer uh or pbsom plus aptainer.
It's it's a a a cifb based um container runtime that runs as a service and accepts work from fuzball orchestrate. And we just made some announcements recently that Fuzball can now also use your existing batch scheduling system as a backend to do its batch scheduling. So instead of actually it it running the queue and doing all those things itself if you already have Slurm or PBS running, you can hook Fuzball into it and you can say okay just make requests on my behalf to you know let users run their workflows um through Slurm or through PBS. Yeah. And in that situation, we're we're kind of using PBS or Slurm like a subresource manager or a provisioner in Fuzball's way of looking at it.
It sees PBS or Slurm as though it's another cloud that it can provision resources out of. And then once it has resources from that system, it then dispatches fuzzball jobs into them as normal. Awesome, gentlemen. Really appreciate this. This has been a fantastic look at service endpoints and a conversation around sovereign AI. Uh, excellent job to you and your team for for all the uh the hard work that went into this. Uh, I know that uh uh I know that it's been uh been a long road, but I mean just just what we've seen over the past couple of weeks with endpoints is absolutely amazing. So, thank you both uh for joining us.
Um, so real quick, Dave, any any closing thoughts? anything you want to send people off onto the road with? >> Yeah. Well, um, just as I was kind of saying about like, um, you know, new technology, just giving you new power, new new ideas and the ability to do stuff that you never even really thought of. I mean, I think that when we started doing service endpoints really, we were just trying to enable HPC engineers who wanted to run things like Jupyter notebooks. >> Yeah. It was all Jupiter. >> Yeah. And now we're like, "Oh my gosh, now we can we can host our own, you know, AI models and we can host them internally within a company or publicly or, you know, whatever." And like it just it just keeps on going.
So yeah, like m this over and think about it and I think we'd be really excited to hear what kinds of things you want to do with this. >> Yeah, absolutely. >> Yeah, it uh local AI was was an really amazing addition to to that list of use cases, but the one that caught me uh was VDI. that that's a really awesome uh use case that a lot of people the especially if if you need some kind of graphical uh uh device attached to your VDI this this is an amazing use case for that. So that's that was that was really cool to see. >> Well that's a super broad set of potential use cases for a what seems like such a small feature and that's part of what makes fuzzball so special in this space.
Lots of things that are targeting AI are just that and maybe even just one piece of software in that stack or or one API or one platform for doing it. Fuzzball is a generalpurpose platform that can use its features to enable a virtual desktop or an AI model. Uh and it's it it's doing a good job of spanning all of these use cases. So, Jonathan, uh, la last comments for you. Any, uh, any closing thoughts you want to share with our audience? >> I'll just add on to what Dave said. We'd love to hear what excites you about this and what you would want to use it for, but then we'd hope to get it in front of you.
We'd love to uh, make Fuzzball available to you. So, let us know and reach out and and we'll we'll get started. >> Awesome. And uh as as we close, if you are curious about Fuzzball, about how to make some of these things happen, uh some of our other products as well, definitely head over to our YouTube channel, uh we're going to be creating tons of content from tech tip videos to interviews like this one uh to additional webinars. Um, and if you have an interesting use case for service endpoint or you're you're tackling some sort of challenge uh with sovereign AI, please head over to uh our LinkedIn or uh Twitter X.
I refuse to call it X. It's Twitter. It will always be Twitter. I will die on that hill. Um, but uh head over to uh find CIQ anywhere on on social media and and send us a message. Let us know what uh what you think you might use service endpoints for. let us know what sovereign AI troubles you're you're tackling. Uh so on behalf of my guests today, Dave and Jonathan, thank you all for joining us. Really appreciate you all tuning in. Uh we know that everyone's really really busy, especially in here we are the end of January already and we're all deep inside of all of the work things that we have to do.
So, you know, giving up an hour of your time to hang out with with with the three of us really really appreciate it. And uh and a special thank you to Wolf Gang who did a lot of the uh prep for the for the demo and the infrastructure that you saw today and Gwen who's been helping me uh produce today's uh webinar. Really appreciate both of you and uh so with that said, thank you all for joining us and we look forward to seeing you all in the next webinar in two weeks. Thank you and we'll see you then. Thanks. See you.
Built for scale. Chosen by the world’s best.
2.75M+
Rocky Linux instances
Being used world wide
90%
Of fortune 100 companies
Use CIQ supported technologies
250k
Avg. monthly downloads
Rocky Linux
Have questions about your infrastructure?
Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.