Why we're building OpenWALDO

Why we're building OpenWALDO

Contributors

Gregory Kurtzer, CEO of CIQ

For more than two and a half decades, I've helped build open source projects and communities that became infrastructure the world depends on: Warewulf, CentOS, Singularity (now Apptainer), and Rocky Linux. Each one started the same way, with a gap that a community was better positioned to fill than any single company. Today we're launching OpenWALDO, an open source project building a shared, community-governed corpus of AI training data, sponsored by CIQ, and I think it's the clearest version yet of that same pattern.

A foundation we can share

Every AI builder needs a broad foundation of training material, and right now, every team is building that foundation on its own. Organizations across the industry are separately collecting, cleaning, organizing, licensing, and maintaining much of the same underlying data, work that gets duplicated instead of shared.

Open source software solved a version of this problem decades ago. When the code, the methods, and the build process are open, anyone can inspect what they're building on, trust it, and improve it for everyone else. AI has had an equivalent for models you can run and adapt, open weights, for a while now, and that's been a real and valuable step forward. What it hasn't had is the same openness for the layer underneath the weights: the training data itself. That's the piece OpenWALDO is built to provide.

A shared, community-driven corpus gives every builder a strong starting point, so they can spend their time on what actually makes their work distinct: architecture, training methods, specialized knowledge, safety, and performance, instead of re-collecting the same base material everyone else already has.

"WALDO," and what's still missing

WALDO stands for Weights, Artifacts, Licenses, Data, and Origins, and in my mind, that's close to the full definition of what a truly open model needs to provide. But there's one more piece that a definition alone can't supply: community. OpenWALDO is both of these things at once, a community of people working on AI together, and a community-managed corpus of training material with the toolchain needed to contribute to it, review it, verify it, and train from it.

How the corpus stays accountable

The OpenWALDO index records where training material came from, what rights were asserted over it, who contributed it, and exactly which bytes belong to each version. Models trained from the corpus can carry an AI Bill of Materials that ties the finished model back to the sources, licenses, and training runs that built it.

The mechanics behind this aren't new. The public index lives in Git, the same tool that has governed code review for open source software for years, so changes are readable and attributable. The larger data objects live in federated storage and are verified by content hashes, so what you download is probably what was reviewed. Contributors sign off under the Developer Certificate of Origin before their work goes in, the same accountability mechanism major open source projects already rely on. We're not inventing new trust infrastructure here. We're applying the infrastructure that already works.

Who this is for

OpenWALDO gets more valuable every time someone contributes knowledge, improves the tools, reviews the evidence, or helps another participant get unstuck.

Researchers and model builders can help define what useful training material actually looks like. Data experts, librarians, and legal reviewers can strengthen licensing evidence and provenance. Infrastructure engineers can help the index and training tools run at scale. Companies can contribute their own public documentation, product information, and knowledge directly, rather than hoping a crawler eventually finds it and gets the details right. And hobbyists and home labbers can start training their own small models on their own hardware, and see what a fully open, provenance-tracked pipeline looks like end to end. While what we have today is not anywhere near a frontier model, it is the beginning, and something we can build upon and scale together.

Labs and model providers, open weight or otherwise, can build on it too. The corpus and its bill of materials work as a verified baseline: take it, add your own proprietary data and your own models on top, and ship with a clear line back to your sources. The foundational work gets done once, in the open, so everyone building on top of it can focus on what makes their model theirs.

Why this has to be community first

AI "source code" doesn't have a strong community of its own yet, not the way Linux does. It has users, and it has labs, but it doesn't have a shared place where the people who care about this can actually build something together. That's what I want OpenWALDO to be, first. The corpus, the community, and any models that come out of it are outcomes of that community, not the point of it.

That's also why I've thought carefully about how CIQ fits in here. I'm the founder of OpenWALDO, and I'll be leading it directly. CIQ, the company I founded, is sponsoring the project, funding the engineering it takes to get this off the ground. Together, the foundation is set, and we open the invitation for others to join us to help build the community behind AI.

Build it with us

For more than two and a half decades, I've watched open source turn ideas into communities and infrastructure that changed entire industries, and every time, it worked because the community brought more to it than any one founder could have alone. I believe OpenWALDO can do the same for AI.

I also like that the name gives us an honest answer to a question people in AI keep asking: where is the training data, where did it come from, and who said you could use it. Here's WALDO. Stripes and all.

Whether you build models, curate data, operate infrastructure, study licensing and provenance, or are just excited to learn how any of this works, there's a place for you here.

Join the community on Slack and GitHub, or read more at openwaldo.org.

Subscribe to our newsletter

Related posts

Ascender Galaxy Proxy is now open source

Ascender Galaxy Proxy is now open source

Ascender Registry: a content hub you actually own

Ascender Registry: a content hub you actually own

Behavioral drift: The risk your eval stack won't catch

Behavioral drift: The risk your eval stack won't catch

The CIQ portal is live: access, evaluate, and deploy CIQ products, on your own terms

The CIQ portal is live: access, evaluate, and deploy CIQ products, on your own terms

Built for scale. Chosen by the world’s best.

2.75M+

Rocky Linux instances

Being used world wide

90%

Of fortune 100 companies

Use CIQ supported technologies

250k

Avg. monthly downloads

Rocky Linux

Have questions about your infrastructure?

Talk to a CIQ engineer about Rocky Linux, HPC, and AI infrastructure.

Talk to an Expert