In the world of artificial intelligence, a war of words is raging. Behind the technical terms “open source” and “open weight” lie issues that are crucial to the future of technology. An explanation of a distinction that will determine who controls the AI of tomorrow.
Artificial intelligence is currently going through a crucial phase of definition. As generative AI models transform our societies, a fundamental question divides industry players: what does “open AI” really mean? Far from being purely semantic, this question determines access to these technologies and their future development.
Two approaches are competing today. On one side are “open weight” models, favoured by many companies. On the other is the authentic “open source” approach championed by free and open-source software organisations. To understand this distinction, we must first understand how an artificial intelligence model works.
The fundamentals: how an AI model works
An artificial intelligence model relies on three essential elements. First, the “weights”: millions or billions of numerical parameters that determine the model’s responses. These weights are obtained through “training”, a process that progressively adjusts these values. Second, the training data: the text, images or other content used to teach the model. Third, the source code: the computer programs that orchestrate the model’s training and operation.
This architecture explains why not all “open” models are created equal. Depending on which elements are shared or kept secret, the possibilities for using and improving them vary considerably.
Open weight: openness with shifting boundaries
The “open weight” approach consists of publishing only the weights of the trained model. This strategy allows developers to use the model and adapt it to their specific needs. However, it keeps the crucial elements of its creation hidden.
In practical terms, receiving an “open weight” model is equivalent to receiving a fully assembled car without the manufacturing plans, the list of components used or the specifications of the production tools. The user can drive the vehicle and even make superficial modifications, but remains unable to understand its internal mechanisms or reproduce its manufacture.
This limitation is not insignificant. Without access to the training data, it is impossible to assess the model’s potential biases or understand its strengths and weaknesses. Without the source code, reproducing the training process becomes impossible, preventing any independent verification of the claimed performance.
Added to this are so-called “open” operating licences that are often more restrictive than existing standards (Apache, MIT or others) and custom-made by model publishers. Meta’s Llama model is a perfect example of these restrictions. Despite being labelled “open”, this model remains unavailable to European users because of legal constraints that the company refuses to lift. This situation reveals the limits of conditional and geographically selective openness.
Authentic open source: the requirement for complete transparency
The Open Source Initiative, the leading organisation in the field of free and open-source software, has established strict criteria for artificial intelligence. A genuinely “open source” model must provide all of its components: the complete weights under an open licence, detailed documentation of the training data, the source code needed to reproduce the training process, and comprehensive technical documentation.
This approach draws on the four fundamental freedoms of free software, adapted to the context of AI. Freedom of use allows the model to be used without restrictions on its application or sector. Freedom to study makes it possible to understand in detail how the model works and how it makes decisions. Freedom to modify allows the model to be adapted to specific needs. Finally, freedom to redistribute encourages improvements to be shared with the entire community.
These principles create a virtuous circle of collaborative innovation. Every improvement can be shared, studied and incorporated by other developers, accelerating technological progress as a whole.
The contrasting landscape of current initiatives
In response to these definitions, industry players are adopting a range of strategies, each with its own benefits and risks.
The pioneers of complete transparency
Organisations such as Eleuther AI, the Allen Institute for AI and HuggingFace have chosen the path of maximum transparency. These projects share not only their models’ weights, but also the training data and creation processes. Their approach makes it possible to reproduce the work in full and independently verify the results.
However, this transparency comes with significant legal risks. Eleuther AI had to remove several components of “The Pile”, its well-known dataset, following copyright challenges. A Dutch development project based on Llama was removed entirely for violating the licence. These incidents reveal the legal grey areas that threaten the open-source ecosystem.
The emergence of legally secure solutions
Faced with these uncertainties, a new generation of initiatives is prioritising legal certainty. The Common Corpus project, for example, compiles exclusively data that can legally be distributed. This approach eliminates copyright risks and allows redistribution without fear of legal action.
Daijobu AI’s models, developed in France, follow a similar philosophy by ensuring compliance with European regulations, including the AI Act and the exceptions provided for text and data mining. Although these models are not necessarily technically “more open”, they provide legal certainty that is crucial to institutional and commercial adoption.
The challenges of licence continuation
Some projects are experimenting with an even stricter approach: “licence continuation”. Under this principle, a model trained on Wikipedia should inherit the encyclopaedia’s licence. Although intellectually coherent, this logic proves virtually unmanageable in practice.
Combining sources under different licences—Creative Commons, the GNU Free Documentation License, the French Open Licence—becomes an insoluble legal puzzle. This approach is viable only for projects based exclusively on the public domain, considerably limiting the possibilities for innovation.
Increasingly open alternatives
DeepSeek’s arrival on the market disrupted the established balance. By releasing its leading models under the fully permissive MIT licence, the Chinese company demonstrated that a radically open approach was not only still possible, but also competitive.
This demonstration exposed the limitations of the partial-openness strategies adopted by other players. When a high-performing model becomes available without restriction, legal subtleties and artificial limitations lose their economic justification.
The impact extends beyond the technical sphere. DeepSeek revealed an uncomfortable reality: many companies exploit the ambiguity between open source and open weight to maximise their profits. They reap the benefits of contributions from the open-source community without genuine reciprocity, while preserving their competitive advantages through the proprietary elements they retain.
The European regulatory framework is taking shape
The European Union is not remaining passive in the face of these issues. The AI Act and the Code of Practice for AI are redefining the rules that apply to artificial intelligence models. In particular, these texts impose mandatory traceability of training data and greater transparency regarding the sources used.
Compliance with the “text and data mining” exception is becoming a legal obligation, not merely good practice. Developers must now document their sources precisely and respect the “opt-out” rights exercised by content holders.
These regulations, perceived by some as constraints, could paradoxically make the market healthier. By imposing clear standards, Europe is forcing industry players to choose between authentic transparency and marketing claims about their supposed “openness”. And it is fostering the emergence of genuinely sovereign AI.
Nevertheless, many uncertainties remain. The use of copyrighted content for training remains controversial, with legal interpretations varying across jurisdictions. This situation discourages innovation and favours organisations with substantial legal resources.
A practical guide for developers
In this complex landscape, developers must adopt a methodical approach when choosing their tools.
For standardised commercial applications, an “open weight” model may be sufficient if there is no need to understand or modify the training processes. This option offers flexibility of use while retaining relative legal simplicity.
By contrast, for research, audits of critical systems or the development of innovative solutions, the complete transparency of open source becomes essential. Only this approach allows a deep understanding of the mechanisms and continuous improvement.
In every case, licences must be examined carefully. Restrictions can be hidden in contractual details, with major implications for the end use. Anticipating regulatory developments by choosing models that already comply with emerging standards is also a wise precaution.
Naturally, Daijobu AI supports you in making these pivotal technology choices for your company’s development.
An issue of technological governance
The distinction between open source and open weight goes far beyond technical considerations. It fundamentally determines who will be able to understand, improve and democratise the technologies transforming our societies.
This battle will define the future balance between open innovation and proprietary control. It directly affects the ability of researchers, public institutions and small businesses to participate in the development of artificial intelligence.
The future is taking shape around two scenarios. The first would see the emergence of a genuinely open ecosystem based on transparency and collaboration. The second would preserve the dominance of a handful of major players using terminological ambiguity to protect their competitive advantages.
