The artificial intelligence sector is currently navigating a profound philosophical and operational schism over the misuse of industry terminology, with prominent technology leaders cautioning that the casual conflation of “open source” and “open weights” threatens to erode decades of hard-won software freedoms. While the terminology of the open-source movement has been enthusiastically adopted by generative AI developers and commercial labs alike, the actual deliverables provided to the developer ecosystem frequently fall significantly short of the foundational tenets historically guaranteed by traditional open-source licensing.
This linguistic blurring came to the forefront during the Open Source Summit Europe in Prague, where Peter Farkas—co-creator of the open-source MongoDB alternative FerretDB and current CEO of open-source database services provider Percona—issued a direct plea to the technology community. Addressing an audience of developers, enterprise architects, and open-source advocates, Farkas emphasized that the integrity of the broader open-source ecosystem is at stake if the AI industry continues to mischaracterize models that offer only partial access to their foundational architecture.
The Core Divide: Access Versus Absolute Transparency
To understand the urgency behind Farkas’s warning, it is essential to examine the technical architecture of modern artificial intelligence systems. An AI model’s weights represent the vast array of numerical parameters generated during the intensive training phase. These weights encode the complex patterns, probabilistic associations, and linguistic structures the model has acquired by processing massive datasets. When a laboratory or enterprise releases an open-weight model, it provides developers with these downloadable numerical parameters. This grants third parties the ability to deploy, fine-tune, and run the model on their proprietary infrastructure, bypassing the need to rely exclusively on managed Application Programming Interfaces (APIs) hosted by the model’s creator.
However, the availability of model weights does not equate to complete transparency. As Farkas pointed out during his Prague keynote, downloading model weights provides only the final output of a complex pipeline, omitting the foundational elements that define genuine open source.
“You don’t have the source code, you don’t have the training data, you only have the output of these two,” Farkas stated. “So why would we call this open source in the first place?”
This perspective is echoed heavily within the academic community. James Landay, director at Stanford University’s Institute for Human-Centered AI (HAI), addressed this precise vulnerability in an extensive analysis published by the institute. Landay established a clear functional dichotomy between the two approaches, characterizing open weights as an answer to a logistical question, whereas open source addresses a systemic, collaborative necessity.
“Open weights answer ‘can I run this?’” Landay wrote. “Open source answers ‘can I trust this, improve it, and build the next thing on top of it?’”
According to Landay and other research analysts, major commercial AI laboratories excel at answering the first question by granting access to weights, but they consistently fail to provide the ingredients necessary to satisfy the second. Without access to the underlying training datasets, data curation methodologies, filtering algorithms, and training source code, independent researchers and enterprise developers cannot independently verify model safety, trace the provenance of generated outputs, or fundamentally rebuild and improve the technology from the ground up.
Market Adoption and the Rise of Open-Weight Dominance
This definitional debate is not merely an academic exercise confined to conference halls and university departments; it is occurring against the backdrop of a massive structural shift in how production AI is consumed globally. Open-weight models have rapidly transitioned from experimental curiosities into foundational pillars of enterprise architectures and developer workflows.
Recent telemetry data highlights the extent of this market penetration. In August, metrics captured by Vercel’s AI Gateway indicated that open-weight models accounted for 56 percent of all token volume processed through the platform. Simultaneously, OpenRouter telemetry revealed that open-weight models captured 60 percent of United States-originating token consumption, with models developed by Chinese entities accounting for the vast majority of that volume.
The widespread adoption of models from developers in China—most notably DeepSeek—has further intensified the debate. DeepSeek’s advanced reasoning models have been widely and frequently described across mainstream media and industry publications as “open source.” Yet, much like their Western counterparts, these releases typically consist of model weights and inference code, while omitting the exhaustive training datasets and pipeline architectures required to reproduce the models independently.
Farkas acknowledges the immense utility that these models provide to the developer community, drawing a careful line between practical utility and semantic accuracy.
“Are open weights a bad thing? No, open weights are great,” Farkas noted. “You can run your models in your own environment, you can experiment with them, and if you understand the risks, you can also use it in production. The problem is when open weights are positioned as, ‘hey, this is as good as open source.’”
The Specter of “Open Washing” and Broader Legal Implications
The risk, according to industry veterans, is that “open washing”—the marketing practice of labeling restricted or partially transparent products with open-source terminology—will become normalized. If the technology sector accepts a diluted, ambiguous definition of open source for artificial intelligence, the damage may not be contained within the machine learning domain.
Farkas cautioned that a degradation of the core definition could ultimately compromise traditional software licensing frameworks. He pointed to potential scenarios where enterprise software vendors might distribute proprietary code under heavily restricted licenses—such as modified versions of the Apache 2.0 license that introduce geographical or operational constraints, such as prohibitions on usage within the European Union—while still marketing the products as open source.
“Open source AI is one thing, but if we talk about open source, and we let this happen to the core definition itself, that is going to go way beyond AI,” Farkas warned.
While the majority of labs maintain restrictive postures regarding their training pipelines, isolated exceptions demonstrate that greater transparency is technically achievable. Xiaomi’s recent launch of its MiMo-V2.6 model family serves as an illustrative case study on the outer edge of the transparency spectrum. In an unusual move, Xiaomi livestreamed nearly a week of reinforcement-learning training via a public dashboard. Furthermore, the company released over 7,000 reinforcement-learning task environments, alongside its reinforcement-learning training code and detailed technical documentation.
While governance purists continue to debate whether even Xiaomi’s comprehensive disclosures satisfy the strictest interpretations of open-source AI, the release underscores the vast operational spectrum that currently exists under the broad umbrella of “open” AI development.
The Evolving Regulatory and Standards Landscape
As the commercial deployment of AI accelerates, establishing a universally accepted standard has become a paramount objective for standard-setting bodies. The Open Source Initiative (OSI) stepped into this contentious arena in 2024 by publishing its inaugural Open Source AI Definition. The framework attempted to establish explicit criteria centered around the fundamental freedoms to use, study, modify, and share AI systems.
However, the OSI’s initial definition immediately encountered fierce debate, particularly regarding the mandatory disclosure requirements surrounding training data. Critics from various sectors argued that the criteria were either too lenient to protect genuine open-source principles or too stringent to accommodate modern commercial and legal realities surrounding copyright and proprietary data.
Addressing Farkas and the broader summit audience, Duane O’Brien—who assumed the role of executive director at the OSI in April—directly acknowledged the validity of these criticisms and announced a concerted effort to recalibrate the organization’s approach.
“We did have a conversation two years ago—it was an important conversation, and two years is a long time in this space,” O’Brien remarked during the Open Source Summit Europe. “We are in the process of reopening and having another set of conversations.”
As part of this consultative expansion, the OSI recently launched its Open Source AI Fellowship program. The initiative appointed Gabriel Toscano as its inaugural fellow to spearhead efforts to build international consensus around the core components of open-source AI. The structured two-year program is designed to evaluate potential revisions to the OSI’s definition through a series of community-driven roundtables and global working sessions, colloquially termed “open source salons.”
O’Brien explicitly invited critics like Farkas to engage directly with the process, signaling a pragmatic willingness within standards organizations to adapt definitions to the rapid evolutionary pace of generative AI technologies.
Conclusion: Navigating the Semantic Future
The ongoing debate over open weights versus open source is far more than a semantic dispute over vocabulary; it is a foundational struggle for control over the future trajectory of computing. As open-weight models capture the majority of token consumption across major developer gateways, the commercial incentive to adopt the halo of open-source credibility without bearing the corresponding transparency burdens remains exceptionally high.
For developers, enterprises, and researchers, maintaining clear distinctions will be vital for risk management, regulatory compliance, and scientific reproducibility. Whether the OSI’s renewed consultative process can successfully forge a unified definition that satisfies both commercial realities and ideological purity remains one of the defining questions for the artificial intelligence industry as it moves deeper into its production era.
