Here's WALDO: AI's first open source project and community

For information on OpenWALDO and the open source community check out: https://openwaldo.org/

Most sovereignty conversations stop at infrastructure: where a model runs, who hosts it, who can see your prompts. That covers three of the four things an organization needs to control. It skips the first and most basic one: the model itself, and whether anyone can verify what actually trained it.

Join Gregory Kurtzer, founder and CEO of CIQ and founder of OpenWALDO and Rocky Linux, and Addison Snell, CEO of Intersect360 Research, for a conversation grounded in two new pieces of research and reporting: Intersect360's enterprise AI sovereignty whitepaper, and the launch of OpenWALDO, an open source project building a shared, community governed corpus of AI training data.

Open weight models solved part of the sovereignty problem: you can download a model and run it on your own terms. But open weight is not open source, and enterprises are increasingly asking the harder question underneath: what data trained this model, under what license, and can I actually check? Kurtzer and Snell will walk through the sovereignty framework enterprises need to apply to AI, where open weights fit and where they fall short, and what a truly open, verifiable foundation for AI looks like in practice.


What we'll cover

  • Why sovereignty has to extend past infrastructure and into the model itself

  • The difference between open weight and open source, and why enterprises should care

  • What an AI Bill of Materials is, and how it lets you verify what trained a model

  • How OpenWALDO's community governed corpus gives labs, enterprises, and researchers a verifiable baseline to build on

  • What the Intersect360 survey data shows about how enterprises are actually thinking about AI control today


Speakers

  • Gregory Kurtzer, CEO of CIQ and Founder of Rocky Linux, CIQ

  • Addison Snell, CEO, Intersect360 Research

Transcript

Welcome everybody and uh you know thanks for joining us today. Uh we're today we're going to be talking about Open Waldo and you know what it means to have greater visibility and control over the AI models that your organization uses. And uh of course today uh joining us is Gregory Curtzer, founder and CEO of CIQ and founder of Open Waldo. We also have uh Addison Snell, CEO of Intersect 360 Research, my boss. My name is Kevin Jackson. I am an analyst here at Intersect 360. And uh Greg and Addison, it's great to have you here. >> It's good to be here. Thank you, Kevin. >> I always have great conversations with Greg, so this is fun for me.

Let's see how it goes today. >> Absolutely. [laughter] >> Great meaning interesting or weird. >> You know, we'll see where we go. >> Exactly. Well, you know, we're going to explore what Open Waldo does, you know, why it matters for enterprises and what it could mean for how organizations evaluate and adopt AI going forward. So, you know, just to get us started, uh, I'm going to put this question to both of you, but Greg, I'd like your response first.

uh you know if I'm running a model on my own infrastructure you know I'm keeping my data private I might feel that I've handled or you know that I've checked the sovereignty box but actually you know what am I still not controlling and why does that matter >> there's a lot of aspects about running your own model trusting your own model trusting the infrastructure that that model is running on and and how do you check that box on on sovereignty um and and It's even more than sovereignty. It's it's ownership. It's control and it's visibility. Knowing what it is doing. Uh so a lot of organizations are pulling openweight models and then running them locally.

We don't know what a lot of these openweight models have been trained on. We don't know what they're what they're going to do uh in a variety of different situations. We don't know if the code that they're going to be writing or if the interfaces and tooling that they're going to be using is um you know, we don't know what they're going to what it's going to be doing with it. I mean, that's kind of it's kind of one of the the the I'll say the benefits of a non-deterministic system such as such as an LLM. But at the same token, even though it's non-deterministic, they definitely have kind of uh directions that they that they will pull towards based on the statistics of of what they were trained on.

View full transcriptHide full transcript

So, you need to know kind of what's inside your model. You need to have proper um uh boundaries, guard rails, and sandboxing of your of your tooling, of your agents, etc. And you need to know what everything is doing. So you need to have proper infrastructure visibility and monitoring. >> Yeah. I mean that's the way we've seen it in our research that of course people are organizations, enterprises are already moving to adopt AI as you know they they want to adopt it. They want the productivity gains. They want the efficiencies. But what I like about what you said, Greg, about control, a lot of that really comes down to risk.

And we've seen this with any new technology that gets adopted, whether it's, you know, high performance computing modeling and simulation or whether it's cloud computing or whether it's the internet, right? What are what are my risks that are involved? And it's not that I have to have no risk. I just want to know what the boundaries are on the risk. Am I protected on the downside on cost? Am I protected on the downside on um on uh you know what uh what catastrophe might happen, right? It's basically the way to put it. And the topic that's in the popular press right now is all around regulation, regulatory environment.

What we find in our surveys is that people feel like I'm kind of okay with understanding what the different regulatory environments are, but I need to know that they're stable, right? And you know, how do I control my intellectual property? How do I control uh my any of my downsides? And that's what having my own model kind of feels like it gives me more of a sense of control. I think, you know, regardless of where you stand on the regulatory debate right now, we need more regulation, we need less regulation. I think what a lot of organizations want is predictability. And I think what a lot of organizations want right now is uh you know, can well I I think the result is can regulation even stay ahead of AI right now?

If I'm going to sit back and wait for regulation to explain everything to me, I'm going to get too far behind. That's just part of the risk environment right now. So, how do I go forward with a model, feel like I have some kind of semblance of control and manage my risk while I chase the productivity gains? >> You you bring up uh such a good point. Um uh a few of them, but but one of them that I want to touch on is the data.

So for an AI model to whe whether you're running it on prem or you're running you know a a a frontier model you need to give it data in order to get a good response right so inside of your prompt or inside of your tooling it has to have information about what you want to be asking it and if you're asking it things about kind of your core business core products trade secrets things under NDA like you have to include bits and pieces of that if not the entire it in its entirety in your prompt in the context, which means you very well might be leaking data.

And this gets really bad when we start talking about agents and um and even more than agents, all of the tools that the agents writes that are scattered all over the infrastructure or even on the internet that have access to your data. And then what's securing those those tooling? what's secure securing those applications and then what's securing the agents to make sure that they are being maintained within their guard rails and that we're not doing something, you know, like hacking another another site or hacking a foreign government. >> Well, well, right. I mean, there are so many analogies with that that I expect my AI to act within a certain boundaries just like anything else that I hire or employ.

If I tell my personal assistant to go pick up my dry cleaning and my personal assistant gets there and finds the dry cleaner to be unexpectedly closed, I don't expect my personal assistant to smash the windows and break in in order to follow the directive of go pick up my dry cleaning, right? Without at least checking with me first about, hey, it's closed. Do you how bad do you need it? Do you want me to Right? That would be outside the normal realm of what I expect. Now, that's where we talk about guard rails and how do models behave. And I guess you can train a model to do whatever you want.

Maybe you're someone who wants it to go break in. Uh but that should be auditable in some way, right? Did you give it that directive or not? And you know, when we design a car, we expect it to behave in certain ways, which doesn't mean you can't intentionally cause damage with a car, but you want it to be built so that you don't accidentally cause damage with a car. I think that's the point. >> That brings up a whole another can of worms, by the way, which is uh accountability and liability. Uh and where does that exist? Kevin, I'm gonna defer to you on that particular [laughter] can of worms to keep us on track.

>> You know, it's it's it's a conversation we're all having right now. I think something that Addison says all the time that I'll fill in here is that if you use a new technology to commit fraud, it's not that, you know, uh we need to write a new law. Fraud is already illegal, right? Uh it's something that Addison brings up quite often. uh we already have laws in place that perhaps we could be using uh and strengthening, but you know, it's a it's a whole debate that could suck up this whole hour if we wanted to. [laughter] >> Well, I I think that's not the point, right?

But but when it comes to control and auditability, there's been a lot of discussion of things like open weights, but then what's the difference between open weight and open source and what do you get out of it? That's and that's really, you know, where I was driving towards, you know, I'm I'm seeing a lot of people use open weight and open source interchangeably and they shouldn't. And, you know, especially Greg, you know, can can you tell me like what's the difference and why does that difference matter to a company a business? Well, the first thing that I want to mention is it's actually it's really confusing and uh it's confusing because if you think I mean several different aspects if you think of um open source software the open source premise was kind of written around the notion of open source software.

So we kind of have a really good firm understanding of what that means when we're applying it to software. You basically think about it as you have the source of something whether that's source code or um you know it's been debated if data is is considered the source in in programs but you have your source and then you have your output and if your source is open source the output is open source kind of that's the definition of open source right so um with with with weights AI weights it's different and it's different because we have a bunch of open- source tooling that we're using to build kind of the the foundation and architecture of our models.

Uh and and then we take the the input data the source which is not open use the open- source tools to create binary weights and these binary weights then are distributed in some cases under an open- source license. So that means you can do whatever you want with the binary weights, but you don't know what went into those binary weights. The source to it is not available. So I'm a purist when it comes to open source in in in the most fundamental way. And as a result of it, I think of it as any any source that you're using to create an artifact needs to be open for that artifact to be called open source.

And so for an AI That means that's your training data. That's the corpus of data you're using. Those it is it is the compose and the build tools that you're using to create your model. And if you just have the open weights, while open weights are great. Please don't get me wrong on this. I love the fact that we have open weights. A lot of the complaints that we have around open weights, a lot of the concerns, the transparency Addison that you were mentioning, well, the the the binary weights are kind of like a black box. You can't ever go from those binary weights and then to understand how are they produced.

And so as a result, we need to kind of be thinking about this from if we want to know what's inside of those and if we want to mitigate the concerns of what it means to be an open model, well, we got to go back to the source and we got to have trust in here's everything that was used to create this model. And if we have that, then we can fundamentally call ourselves an open- source model. But without that, open weight models are not bad, but they're definitely not open source. >> Yeah. I think what you're getting at, Greg, is it hits right at that question that we had about mitigating risk because the or the large organizations or even small organizations we talked to generally have processes in place for when something goes wrong, they want to trace back to how did that happen?

Whether it's a product defect or a failure in policy somehow. and you know there was some customer harm. You want to know not just how do we close this gap but what caused that to happen? What allowed that to happen? And with AI and some of these openw weight models but they're you don't know how they were trained. There's that worry that I'm going to trace it back and I'm going to get back one or two steps and the answer is going to be well the AI did it and it stops there. Well, why did the AI do that? And I'd kind of like to know why the AI did what it did so that I know that I'm not going to repeat the same mistake, which isn't necessarily really fear driven.

It's just kind of best practices in any organization to be able to audit why we're doing the things we're doing, particularly in areas that are continue considered to be strategically important. When we pulled organizations, enterprises, medium and up to large enterprises on where they've adopted AI, we find it really prevalent in areas like coding and office productivity and you know that kind of thing. But it's it's least prevalent in areas having to do with product development and engineering and R&D, which is supposed to be the promise of AI, right? It's going to help us predict the weather and find new products and cure diseases and all these areas that we're used to associating with high performance computing and scientific discovery.

That's where AI is least uh deployed right now. It's where the biggest gap is. And I think this notion of control and auditability is is one of those pieces that's been missing. You know, there there's multiple science fiction books and and movies and whatnot written about like we we have we have put some of our content, some of our what we're what we're excited about in in humanity and we've sent it up into the stars and aliens have received that and then they have assumed that this is what our civilization is like. I don't know why, but the movie Pixels, I think that's the name of it, with Adam Sandler is popping up in my head right now where they did that with some video games and then the aliens kind of figured, well, this is how humans um operate.

This is how they deal. And so, they had to go beat a bunch of real life kind of video game stuff and whatnot. Um, metaphorically, it kind of reminds me of what we're doing with AI a little bit. We're taking all this training data, which is a gigantic corpus of of of knowledge within humanity, and in most cases, we're just trying to get um uh the the the quantity, not the quality. And in that quantity, well, you know, there's quite a bit of science fiction about how AI has completely destroyed humanity and how we are at war with AI.

Like just as a like if you if you know how how training a model works, you'll recognize immediately like there's tons of weights and how everything statistically kind of maps parameters of tokens and and comes up with kind of the next thing that it's going to come up with and what it's thinking about etc. thinking about uh etc. Um, but if you are training it on things that uh kind of declare war between AI and humanity, I mean, don't be surprised if you ask the AI then like uh, you know, should I be concerned that you're going to kill us all? And the AI is like going back to the Terminator series going, "Yep." >> I mean, look, I'm not a doomer around this.

I think there there will be things that happen that are bad. Just like with any new technology, if you vent a car, you're going to have car crashes. O overall net the car accelerates a lot of things and and does a lot of good for humanity, right? And and you could say that about any new technology. Overall, I think organizations do want to chase the productivity gains here. People are adopting it. The question I keep coming back to is then how do I mitigate the risk? How do I un control understand the downside and in order to activate the upside because AI should be accelerating discovery in all of these ways.

It just feels like there's been a piece missing to help get to that level of control that gives me confidence to move it into those areas of of forwardleaning you know R&D and scientific discovery. >> Absolutely. And you know let's talk about that you know missing piece let's make it a little bit more concrete with this discussion you know discussing open Waldo what could an organization do with open Waldo today uh and what decision would it help that organization make like what what how is this going to further that organization's mission >> so you start by saying what open waldo is we just dropped this term out here >> yeah please Greg why don't you give us a good deep uh discussion of the the entirety of this >> yeah absolutely Absolutely.

So, kind of going back to the open source versus the open weights discussion, um there isn't really a very, you know, uh there isn't a a way for for community members to kind of join together and band together to start creating an AI from scratch. Um is this would involve, as I mentioned before, this involves the AI corpus, so training material. This involves data um that you're going to be, you know, potentially post- training on and and teaching it. For example, how to how to interact with with an API um how to interact with a data structure, uh how to actually, you know, respond properly. So, all of these things have to go into the training data uh into the corpus and then you have your tooling and then all the tooling kind of brings this whole thing together.

Uh a and then you of course need a community around this to to be able to say well you know I'm going to bring in this piece of data or I'm going to work on this tooling or I'm going to be reviewing this corpus uh or I'm going to be looking at the the different types of compose files that we have for creating different types of AI models and I want to ensure and again I use this I use this metaphor tongue and cheek humorously but just to make sure like there there's no um you know Terminator uh uh movie scripts in it or or books and whatnot on um the AI being not aligned with with hum humanity and laws, right?

So, let's make sure that we are putting laws into this. Let's make sure that we are putting appropriate ethics into this. And having that transparency and visibility, Addison, that you're that you're mentioning is super important because if I have the ability to choose between two types of models and I know what went into both of them, I can make an educated decision on which model do I want to run. and I can make that decision as a company now managing some of my risks and and compliance. >> I I think that's exactly the point that we were building up to, right? You've created an AI bill of materials that I can investigate and that the community as it will, you know, if you think about the open-source community and this becomes an an AI community that's contributing to this corpus.

If you've got a question about it, take a look. Here it is under the covers. And boy, when I think back to how open-source software really took off, um, you know, going back to the advent of of uh, Linux in the mid to late 1990s, I that's exactly what happened is you had this terrific open-source community of programmers that were then all working on something together. And it didn't take very long before that raced to the forefront of this is what we're going to build our organization on. >> Yeah. Actually going back to origins is a really great way of looking at this. So um the entire GNU movement started with Stallman who wanted to change the firmware in a printer because he was having issues with h how the printer was behaving.

And um I I I believe I'm I'm not I was not there. So, I can't say for sure, but I believe as I understand that um it wasn't a difficult fix. And he tried even working with the companies to try to fix uh how the printer was behaving and what the printer was doing. But because the printer was this black box and you didn't you couldn't figure out what was inside, he had a really hard time doing that until he found the was able to get the source. And that's what really started the entire free software and open- source movement is if we all have the ability to look at the source for whatever artifacts we're talking about, right?

Whether that source and again I'm a purist, so whether that source is software and code or whether that's whatever went into the artifact itself. If you have that ability to see into it, you now have a level of trust that doesn't doesn't require faith. It required because you could see it. >> Absolutely. And you know this kind of brings me to something I wanted to discuss here. You know, most companies, most enterprises, most businesses, they're not training foundation models themselves, right? You know, and and if I'm using or maybe fine-tuning someone else's model, how can Open Waldo help me? And you know, what would it take to use it in order to benefit my organization?

So training a model I want to say is a very difficult thing to do like there is a lot of things that can go wrong and it is exceedingly expensive to have the appropriate hardware to train a model and an interesting aspect about it is and I'm just going to kind of talk about this very superficially but um training a model and running your model are two different things right so you have your training side which is where you're taking all that input data. You're taking your your compose or however you're kind of managing the the orchestration of how you're building that model and you're running it on a very large very expensive resource usually with many many GPUs.

Um and and and on the light side this is probably tens of millions of dollars to build one of these systems. Um the the large frontier uh AI labs have ridiculously large investments into their training infrastructure. And I want to be very clear like you could be using that training that that that infrastructure and after months of training realize that the model went a skew or your training um uh plan went a skew and you're either starting over or starting over from the last known good checkpoint. And it's an expensive, very frustrating, very high risk kind of process to train a model. And a lot of times you end up with a model that comes out the other side that is not what you expected and not meeting the benchmark of what you needed to do.

We've all, I think, seen this if you've been using AI for any amount of time where it's like a newer release of a model within a lineage or a series feels worse than the previous version or or it's doing something kind of different than you might have been expecting. Um, and that's all because the training process is very hard. So, but once you've trained it, you now have these binary files, these weights, and whether you're talking open weight, open source, or even frontier labs, right? You've got your model and that model is fixed. That model's not going to change anymore. As you start doing what's called inferencing.

So you've got your training on one side, you got your inferencing on the other and you're taking that fixed model, that fixed brain in a matter of speaking and you're going to now go and allow people to use it. Now the interesting thing and a lot of discussions have come up on this which are very tangental to this but um just to quickly touch on it. The reason why AI models appear to learn, they appear to grow, and they appear to understand what you're doing is because of the context window. So as you're doing AI, as you're as you're querying the AI, you have this data structure that the agents and whatnot can pass to it.

The AI model does not change for you. All right? So it is a fixed model. So when we're talking about AI, there's the training side, which as of right now, even with openweight models, is a black box. We don't know what's happening there, right? The inferencing side is where you're running that model. And we're not going to we're not going to go into unpack that a whole lot for this webinar, but that'd be a fun next webinar maybe. But on the training side, this is a process which we've actually been doing and and Addison alluded to this already a couple times um to HPC. The training side of AI is an a it's an HPC problem.

Um and we use HPC tooling in order to build um those models even down to the type of scheduler and architecture that we're using to build those systems. This is 30 plus years of knowledge of how to build the infrastructure to scale up and train AI um is all kind of built within that. Um, now now that I've said all that and laid some groundwork and talked way too much, Kevin, what the heck was your question again? [laughter] >> It doesn't matter. You know, you did a great job of just discussing, you know, I was saying just like most enterprise uh they're not like training their own and if I'm fine-tuning someone else's model, how can Open Waldo help me?

>> Yeah, the question was around how does Open Waldo help? And I think you laid the groundwork for that, Greg, because >> you know, I don't really want to train a whole foundation model myself. I'd like, let's take something simple like a chatbot. >> I want to use that for my business, whether it's internal or for my customers. I don't need to retrain basic human vocabulary from scr from scratch.

I do have to give it some knowledge about my products and my business that my own internal jargon that's important to me that might not be in the general corpus right um so there is some training but I'd like to start from an auditable understood foundation because I don't want to find you know later that my chatbot has been you know picked up a habit of sexually harassing the people who interact with the chatbot because it learned from the internet, right? That would be bad. I don't want that to happen. I don't want to train that part. I'd like to have something that comes in starting out good and then I'm going to specialize it.

>> Actually, that brings up and now I'm doing Kevin's job. Sorry, Kevin. There's a question that I just saw come through. Would you mind hitting to hitting on that question? >> Yeah, absolutely. You know, I uh see this is uh from someone I know, Douglas Edline, great uh colleague. Uh so, you know, the alignment issue uh that he's discussing here. He wants to who determines the alignment. You know, he he says your alignment is fine as long as I agree with it because I'm most I'm never wrong mostly. [laughter] >> I know Doug, too. So, I'm not going to argue with that statement. Um, with that with that being said, um, >> well, I mean, this comes down to training again, right?

We use the term training like you talk about training a dog. And the metaphor I use sometimes is if I train my dog to bite people and it does as far as the dog is concerned, that was the right behavior. Is it the right behavior? I don't know. That kind of depends. Is it a police dog? Was it stopping a crime? Or did it bite my son's friend at a birthday party? Right? But the dog doesn't know, right? The dog does what it's trained to do, which, you know, is a long way around of saying it's a new technology and people will use the new technology like they're going to use it.

And if there's bad actors, there's going to be bad actors. If I invent a car, someone will commit a crime using a car. If I invent a computer, someone will commit a crime using a computer. I don't think you can regulate or wish away bad actors. I do think we should talk about what are the remedies and uh you know in terms of you know how what where's the kill switch? How do I prevent a a bad actor from harming me too much? That's kind of beyond the scope of this conversation. And I think what this conversation is more focused on is how do I keep keep my tool from becoming a bad actor by accident, right?

And that's where Open Waldo has a big role. >> Yep. >> I was trying to set you up there, Greg. All you did was agree with me. So, [laughter] >> fine. I win. >> That was perfect, though. I totally >> No, no, it's it's a great discussion. into something uh that I'm sure we'll touch back on, but you know, just just to keep this rolling. You know, we we've kind of been circling around this, but I'd like to dive into it specifically. You know, say I'm a decision maker at my company and I'm deciding on, you know, how I want to move forward with a project.

You know, how should I decide whether to buy a supported AI product or build on open source? You know, what what should drive that choice? So, um, I'll jump in on that one. Uh, I I I would I would make the relation back to what we've been doing in open source, um, for for decades now. Um, there are companies that stand behind products that will give you guarantees that will give you, um, uh, asurances and, uh, and take the liability on in many cases for those products to at least do what those products say that they're going to do to a reasonable extent. And um uh when when when organizations are buying part of their software infrastructure, that's part of what they're buying.

Even if it is 100% based on open source software, you're buying that that commercial relationship. Um in in in a more direct sense, you're buying the throat to choke. Uh and you are um uh mitigating your your risks and compliance needs by doing so. Now, if a company wishes to go direct to open source and download an open source operating system or build their own kernel or, you know, run without a support contract, without a uh commercial agreement, what they're basically ending up with is well, they're taking on the liability themselves. Now, with open source, it's not an unreasonable thing in many cases to take on that liability yourself because you can see what's in it.

You know what that software is going to do and you can audit if you're so either inclined and or able. You can audit that code and make sure that it's at least as good as as as you know it it it is supposed to be with what we're seeing right now with open weights. And again I want to be very clear I very much support open weights and I think it is a fantastic progression towards open source for this. Uh so I really support open weights. Um, but if you download an open weight uh model from the internet from hugging face and you go and run that in production, who owns the liability of that?

It's 100% on the person who is or the company who has downloaded and running that. And it's a little scary to me when it's a black box and you don't know what's actually in it. Now, I is there going to be compliance and and risk mitigation for that? Probably. But as of right now um it does scare me. So if I have the option and the best case scenario there's there are companies doing this. Uh one such company is RCAI and RC has an openweight model American company openweight models. Matter of fact I think they're the leading openweight model in in the US and they have an openweight model.

You can go to Hugging Face and download it for free or you can go work with them as a company and you can now have the asurances and and um mitigations for controls and and liability and whatnot that you'd have with any open-source provider of software. So, what would be my recommendation? Well, as a company, I'd want a company uh another company to take on that liability and risk as much as possible. I think Kevin, just the fact that you're asking the question underlines the point that there's going to be segments for all of these things, just like there has been ever since Linux and the early debates in the late 1990s about whether to trust my enterprise on open-source software.

I think we found over time that a lot of people did or what I want is something that's open source, but I want a supported version of the open source, right? But either way, taking the boom in AI now and all of the dialogue we're already having around open weights and saying, "What about something more like open source with an with an open corpus and a communitydriven effort?" I mean, I think this is something that really is going to take off. Open Waldo is kind of an early mover here in first to market has the opporto in the future. I guess that's for CIQ and to you know to try to drive their product but we we haven't seen this before and I think to for those who are worried about you know uh US or Western technology and what's going on with the Chinese.

I mean you know imagine if this were a Chinese product right now. Would there be concern about what's the control around it and who's auditing it? I I think the idea that you have an open community, open corpus, open weight, open source, like here's something I can understand and build AI on it. Companies will evaluate how much control, how much support they want over time. But an open community effort around AI, I think we've seen enough in the past to know that this is an idea that ought to take off. >> You know, we've we've been we introduced Open Waldo. We talked about it and um and you just kind of gave this this string of openness that I think is really important.

And it's funny because I gloss over this so often because I maybe I'm so close to it, but the name itself is actually really cute. Not because Waldo is cute. It's because it, well, maybe, but it's because that uh Waldo stands for open weights, open artifacts, open licenses, open data, and open origins. And when you put that all together, that's basically what we're ending up here in terms of what does it mean to be a truly open-source model. And when we're talking about the corpus, we're talking about the data and the origins of that model. So, um, anyway, just wanted to put that out there that all those things together is really where Open Waldo got its name.

>> And it's a it's a really interesting, uh, technology. Me and Addison obviously have been able to speak with you about it. I'm glad we're able to share. Uh, just a reminder to anybody watching, please send us some questions if you'd uh, if you're interested. I've of course got a litany of them, but we'd love to hear from you. But uh you know we're just talking about transparency and I'm interested in something. You know making the training data inspectable is one thing but having the expertise to evaluate it is another. You know how does an enterprise turn that transparency into confidence and you know who should do the checking.

>> I'll take another stab at this. um not to dominate uh the and hog the mic but um again I go back to the how open source works today like uh we holistically as a community as of of enterprises and individuals and people we all rely on open-source software every single day whether you know it or not and there are I mean I can't even tell you how many millions and billions of lines of software that we are relying on that is open source. Uh we have several ways of of trusting that because there are there are examples of uh malicious actors, bad actors trying to infiltrate open source communities to try to get uh malicious code into the open-source ecosystem.

And this has happened 100%. And uh we've we've learned lessons as a result of that. Are the lessons foolproof? Um no, they're not. Uh but we're continuously learning. And if you look at the amount of contributions that we get, it's an incredibly small amount of contributions that have actually turned out to be malic malicious. And what we've been able to do is leveraging things like the developer certificate of origin DCO leveraging accountability um understanding who the contributors are and being able to monitor those contributions with um uh reviews code reviews data reviews etc. uh before we allow anything into the corpus of whether it's open source software or the openwaldo corpus um that creates trust.

The fact that you can even though many people might not be able to either technically or it may just be too much data. The Linux kernel has tens of millions of I think tens of millions of lines. It's definitely in the millions. I don't know if it's tens of millions. Um need to go back and check that now. Um but there's so much code in there and it requires such um specific expertise that even many software developers who are very senior in their roles cannot audit the Linux kernel. They can't audit uh the core tool chains of the C library and the compilers and whatnot within every single uh operating system that we're using that's based on open source.

Can't do it. So the but but specifically knowing that others can is where you start to develop confidence and trust. And when we're talking about data like what's in Waldo, we're not talking about the the the the goalposts are a little bit shifted because we're not talking about um software and code that that people can anybody can review. In this case, we're we're actually just talking about such a mass quantity of data, it's hard to review. We're talking about hundreds of billions and trillions of of tokens going into an AI corpus. U for a small model, you you could be talking about um hundreds of millions.

Um for a large model, you're talking about trillions. And that's a lot of data to review. But if you know what if you have metadata on all of your on all of this corpus and you know each shard what that shard is is responsible for bringing in. It gives you the ability to start pinpointing down and to start saying well here's the whole corpus but here's what I'm seeing in the AI. I want to actually fine-tune and get into a very specific section of that corpus. And even though it's still going to be very very large it's now a reasonable task. And inside each shard of data um which is which is contained within a parquet file, you have different rows.

You can think of it almost like a database in a file. And you have these different rows inside of there. And each row is queryable and gives you the ability to say here's even more metadata, more specific metadata about the content that that includes. And so you can start really narrowing down and kind of following through the hierarchy of the data to get to the point where yeah you can actually you can do this. Yeah, I think you really highlighted uh important things there, Greg, with regard to how these things get adopted. And I would start by looking at what are the processes we already have in place.

Again, not reinventing things, but then we'll find that processes have to evolve somehow. My I like that you were talking about software development in general. Vibe coding is one of these areas that's really taken off and there's this idea of look how much faster I can produce software. Well, but any professional software engineer will tell you that writing the code is one part of it. You know, have I checked it for dependencies against all the other modules? Is it well documented? How about security checking all that software? Right? Who's doing all that as you pile it on? Is it another AI that's doing the security check?

What's that process like? There a lot goes into software engineering beyond just hey I produced all this code that needs to be checked. It needs to be you know integrated in certain ways and there should be a process for that. Maybe the process needs to evolve because you're producing software so much faster. And I can extend that to scientific discovery and you know, hey, I'm producing papers so much faster now because AI is helping with the research. Okay, what about peer review, right? I've just put out 2,000 papers. Who's doing the peer review on 2,00 papers? Is peer review important? it's been a cornerstone of our scientific process.

Maybe you're saying we don't need it anymore. I think probably what you're saying is it needs to evolve. We need to figure out how does peer review evolve to keep up. And you know ultimately I think we learn a lot from our heritage in how we do science and engineering and how we've adopted to other technologies over time. I visited a company once that was starting to use modeling and simulation in its new product designs to help get time to market and and more responsiveness faster. The products once designed still had to go through the same testing they would go through before like all right, I'm going to put it in this stressor in this bender and I'm going to press um you know to make sure it still meets spec.

But the process had to evolve in a way because the old way that the physical guys were doing it is they would press until they knew it was in spec and then stop. And the guy who was designing the modeling and simulation piece was begging them, please, I want you to keep pressing all the way to break regardless of whether that's past inspect because I would like to see does the break match what I predicted in the simulation. Right? because I'm trying to build faith in the simulation. Yeah, I know it's inspect. Keep going. I want that data, right? And it was like this change in policy of how they want to do their testing, which implied some extra cost, but it was a worthwhile cost in order to get the gain.

Now, AI is going to be a little different, but it's kind of the same. What I want is to start with my existing processes and figure out how do I adapt the policies to get the AI gained but still keep confidence and in fact build confidence that the AI is taking me in the right direction and again having an auditable or or open AI that I can look at and say this is how we got to where we are. That's what we already went through with modeling and simulation, right? How do we build the confidence that the simulation is good is it rhymes with the idea of how do I build confidence that the AI is good means evolving my process so I feel confident in the peer review?

I feel confident in the product test. I feel confident in the customer interaction. You know, I I if I have a customer service rep, I trained them and I monitored them to make sure that they're interacting with the customer the right way. Chatbot, same thing, right? I want to train them. I want to monitor them, but the process has to evolve to incorporate the AI if I'm going to really take advantage of it. [snorts] >> You know, I'm glad you mentioned confidence, uh, Addison, because we have another audience question that kind of dives into this. you know, this this this audience member uh seems to run into issues where they they ask the question, where did that particular answer come from, right?

And they they say that open Waldo seems to be a great solution for this, you know, Greg, how far can open Waldo take us toward that today? And you know, maybe what's the difference between tracing a model's training data and tracing an individual response? >> I think I think both are incredibly important. Um uh the way the way models work is you may get something that is verbatim that you can go go through and trace back to an actual document. Um but statistically you might not. And as a result of that it may be pulling in many different documents together to give you one kind of concise summarized answer.

and um and being able to trace through here's what the AI said to where's the data is much much easier when you actually have the data and so the two go hand inand uh right now open Waldo does not contain facilities to trace there are other systems out there so um OMO trace for example exists and um and and over time there's this is just going to continue getting better and better. But I want to really reiterate that you can't always resolve a an answer from an AI to a single like here's where that came from. It's most probably going to come from hundreds if not thousands or hundreds of thousands of documents and tokens and parameters that are all kind of statistically giving you an answer or leading you leading the answer to a particular direction.

So it's AI models are stat statistical masterpieces of how they work and they're very they're they're non-deterministic. So uh you as many people probably found you can ask a model the same question many times and get different different responses but you can also ask a model the same question and sometimes get the exact same responses. Um, I was building an AI agent a while ago. Um, and one of the questions that I would just say is tell me a joke and I found that it would tell me the same joke several times in a row and then sometimes so in the agent I actually put in something that basically said, you know, if you get the same response, don't accept it.

send it back over to the AI and give it information that you got the same response and you shouldn't get the same response. Even if I'm asking the same stupid question over and over, I want the AI to be able to to reiterate and think about it or give me a different perspective or um reconsider and um yeah, so even though it's non-deterministic, well, the thing about a non-deterministic system is sometimes you get the same answer multiple times. So anyway, but sometimes you don't. I mean, I I love this answer, Greg, and I'll just take it and make it personal because, you know, here we are.

We're a couple of talking heads on a webinar, and I love talking to you, and I've talked to you at conferences in the past. We've got a good rapport and you and I are capable of talking about an awful lot of things. We were three minutes to go before this webinar. True story. We were talking about karaoke right before go live. And the people behind the scenes are freaking out because they think the webinar is going to go live. We're going to be talking about karaoke. Did you and I get any advanced training before this webinar about, hey, please you two stay on topic? We did.

People want us to be on topic for this webinar. We get training, too. And I like to think that as a couple of autonomous, intelligent people, we're capable of having a sense of occasion and here's what we're talking about. But then also the other part of that is where did the answer come from? And this is something that's personally very important to me as a professional analyst, right? Because you, Greg, have challenged me on where did that insight come from, especially on things that are not common knowledge, right? How did you reach that conclusion? And I'm able to show you here's the survey, right? Or here's how the methodology, here's the way we built that, here's where we got the list, here's the other corroborating evidence.

Kevin has seen it and you know heard us talk about when we go and measure a new thing we try to do it with three different methodologies to see if we can get the same answer three different ways. This is really important to me as an analyst because that's going to be our niche and I think that's a niche that s potentially survives in the a era of AI. If the reason you want to work with an analyst is just to justify the decision that you've already made, I'll tell you what, you don't need to work with us for that. You can go to any AI agent that you like and say, "Give me a justification for this." And it will do it for you.

And you'll say, "Hey, I asked the AI and it told me to do it." You don't need an analyst for common knowledge. I think you do need an analyst who's going to give repeated different perspectives. If what you want is uncommon knowledge, who says, "I've looked at this and and I think there's a lane here that's underappreciated and this is important to our line of business. I think it's important to your line of business." And when I talked about the perspective of enterprises and control and where are these things deployed? There's a lot of careful methodology that went into that that you can audit that we can audit.

We can take you through it. This is important. And I, you know, the reason that you work with a company like us or with a company like yours is to get that kind of thoughtful leadership. It's an important question and it's one that's important to me personally and I I kind of appreciate it the chance to give a personal answer. >> Yeah, absolutely. So, you know, running low on time obviously. I know you two we could make this a six-hour talk and we'd still have be >> we might start talking about karaoke. I don't know.

[laughter] But you know I I I would like to get a couple more questions and one uh you know looking forward I I I hope to see a beautiful future for open Waldo but what has to happen for open Waldo to become something enterprises rely on rather than an interesting open- source project that they may they watch from the sidelines you know what are the the real nitty-gritty solutions this is going to give >> I'll go first on that just because Greg has been hogging things but also because it's my job to give industry context text, right? And we already laid out what the state of the market is with AI, right?

And the state of the market is A, I want to adopt AI and I'm spending money on it. I think that's a chipshot. People are there and that part is common knowledge, but b there's a gap still between how I'm using AI and what all the promises of AI are right now. And that is related to C this very live debate around AI risks and regulations and accountability and how do I protect the downside there and yeah what is the difference between openw weight and opensource and how do I have something that's auditable that I can move forward those are the clear guidelines that we see in the market right now is openwaldo the answer it certainly seems to me like It's an answer to a lot of these things.

And now I'll pass it over to Greg because you can start distinguishing your product from any others that are out there. But that's the industry the way we at Intersect 360 Research see it today. Um I'm going to approach this very different perspective um uh and focus more on the challenge uh with open Waldo and where um where it is today uh and and how people can and can be part of this. So so people aren't standing on the sidelines. People are are are learning. So the way that I see it is we have people on on one side who totally understand open source. They they understand the value.

They they rely on the value of it. They've been running software infrastructures at large organizations, enterprises, etc. or or hobbies and home labs for decades. They get open source. And we got on the other side. We have a lot of people that have been uh really driving the research and science of AI and building labs, building capabilities. And um it's it's a fairly small population of people um but it's growing and it's growing quickly. So we have the open source people and we have the AI people and there's not much overlap between the two and open Waldo sits at the intersection point of those two communities.

So on one side I'm talking to the open source people which is my background my legacy it's been on that side. So I'm talking with the open source people on the value of an open-source infrastructure the same that we've been leveraging and relying on for decades but now also for this new amazing thing that's going to profoundly change how everybody does their job works and and potentially lives in a good way. I'm not a doomer at all in a good way. Um, and then on the other side, I'm learning as much as I can about AI. Um, obviously I have to as I am running this open source project around AI, I'm learning about building models.

I'm learning about the math and the statistics that needs to go into it and and whatnot. But I'm also communicating with a lot of the people on that side now educating them on the need for open source and what how this will make things better uh for them. The way open source works is it levels everybody up together and it becomes a resource that even direct competitors can leverage and work with and collaborate on. So they're not they're not collaborating on excuse me they're not they're not they're not competing on the general knowledge they're competing on their value ad and it gives them a and again levels everybody up and gives everybody a better opportunity uh to to focus on what they do best.

So Open Waldo kind of drives this and of course I'm getting a call right now there. So, Open Waldo drives us. If this is one of the people that's listening on the webinar, not funny. Um, [laughter] so, uh, Open Waldo really drives us by building a community and building a capability for everybody to come together and and just to give an example of one area that I think that this might be super impactful to businesses, not only on we're going to go train and run our own models or fine-tune our models, which is a huge potential use case. We kind of briefly touched on it, but training an existing openweight model with your specific corpus of data is incredibly valuable and open Waldo provides tooling to make that much much simpler.

Um, but if you look at it from a whole another side, which is companies that may or may not be um even using AI, but uh the example that I always give, and sometimes people don't like the example, but I love the example and I got the mic right now. So, I'm going to talk about it, which is um uh for for decades now, companies have been thinking of marketing uh SEO and search engine optimization. >> And uh one potential like gotcha with search engine optimization is the algorithms are always changing by the search engines. You're always changing your website hoping that something gets slurped up by a a web crawler, gets indexed and cached such that when somebody does a a search, your site shows up and ranks.

Well, in AI, we're approaching it almost the exact same way where companies are just putting a load of data on the internet and hoping, praying that that data gets slurped up by some AI indexer and gets included in the you have no ability to say, "Here's what I want." You're basically just saying, "I'm going to throw a whole bunch of stuff up there and and hope for the best." And it's it's a difficult situation. But what if those companies can just say here's our entire corpus of everything we have for the company um and uh all of our product information, all of our company information, the services that we offer, why we are great, etc., all their marketing material, and they put it directly into the corpus.

And then every AI that slurps up from that corpus now has information about every company. And that information is auditable. And one thing I didn't mention is everything in the corpus, there's one requirement. It has to be distributable. You can't put copyrighted material into the corpus. It has to be distributable, right? So you have to be able to um allow others to take that same data and go and do what they want with it. So incredibly important notion of of open source in general. And this gives us a platform to do this together. And now the comp all these companies can come together and start saying well I want to make sure I've got more information on my product than our competitors.

So they're going to start building more in the corpus and providing more information in that corpus about them about their and all of a sudden the problem of data doesn't go away but it surely gets better. Now all of a sudden we've got really good data in that corpus that companies can be using to train and AI labs be using to train and audit and verify. And I'm going to now put the mic down. Not really a drop. Just very gently put it down because >> I have nothing to add. That was perfect. [laughter] >> This is what I'm saying. This could be a sixh hour uh webinar maybe next time.

Uh but obviously a lot of great information. We're we're running out of time here. So uh >> Oh, I have a plug. I have a plug before you close. >> Please go right ahead. So Addison and I are going to be sitting sitting down together again at supercomputing this year. If you're not familiar with supercomputing, you need to check out supercomputing. It is one of the coolest uh uh conferences in existence. And Addison and I are going to be sitting down and doing a panel. We're going to be having discussions like this. We're going to be taking more information, more questions from the crowd. And uh we really are a lot of fun on stage together.

I'm just saying. So you you you all should come and join. If you're headed to supercomputing, definitely join us. If you're not going to supercomputing, you're missing out. You should >> Craig's going to prompt me to tell a joke while we're up there. I'm pretty sure. [laughter] >> I'm gonna have a couple of dad jokes at the ready. [laughter] Absolutely. >> Uh, yes. >> All right. Yeah. I'm supposed to ask or someone is where do people learn more about Open Waldo? What do you want people where do you want people to go? >> Easiest way is openwaldo.org. No hyphens, no spaces. So, I don't think you need to say spaces anymore.

It's a URL, but no hyphens, no dots, okay, sorry for theorg, but anyway, openwaldo.org, you'll find it. Um, come and join. We've got a Slack instance, super friendly people. Uh, we're all talking about AI. We're talking about building AI and you can learn, you can you can level up and be part of the cool kids club. >> And and don't we all want to be part of the cool kids club? >> We're the commercial for the cool kids. [laughter] >> Oh. Actually, that's probably okay. Well, you guys are cooler than me, so >> Well, uh again, I'll just say that that uh website is openwaldo.org. If anybody has any uh questions, concerns, comments, please head there.

Uh there is a contact us uh section. Greg, Addison, thank you both. We covered a lot here today and uh I think this was a real interesting discussion. >> Thanks, Kevin. Thanks, Greg. >> Thank you, Addison. Bye everyone.

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

9

Enterprise products

Spanning the kernel to the orchestrator

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert