# A Journey toward defining Open Source AI: presentation at Open Source Summit Europe

**URL:** <https://discuss.opensource.org/t/a-journey-toward-defining-open-source-ai-presentation-at-open-source-summit-europe/624>\
**Category:** Open Source AI\
**Created:** [October 1, 2024, 5:05pm UTC](https://discuss.opensource.org/t/a-journey-toward-defining-open-source-ai-presentation-at-open-source-summit-europe/624 "2024-10-01T17:05:58Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![anon18632855](https://avatars.discourse-cdn.com/v4/letter/a/dfb087/32.png) [@anon18632855](https://discuss.opensource.org/u/anon18632855)\
**Post date:** [October 1, 2024, 8:17pm UTC](https://discuss.opensource.org/t/a-journey-toward-defining-open-source-ai-presentation-at-open-source-summit-europe/624/2 "2024-10-01T20:17:17Z")

</div>

> [@stefano](#):
>
> Code and weights need to be covered by an OSI-approved license because they represent the modifiable core of AI systems. However, data doesn’t meet the same criteria.

Right, practitioners are not typically modifying the data in its system of record, rather transforming (selecting, filtering, de-duplicating, segmenting, tokenising, balancing, normalising, etc.) the “source” before “compiling” it to produce a “binary” (model).

> [@stefano](#):
>
> Instead, we concluded that while data is essential for understanding and studying the system[…]

How? By voting? By relying on the MOF? Consensus? Coin toss? It doesn’t matter because you accept “data is essential for […] _studying_ the system” — one of the four _essential_ freedoms — so data **must** be required by OSAID.

> [@stefano](#):
>
> it’s not the “preferred form” for making modifications.

It is the **only** form for making many/most modifications, so it **must** be the preferred form, or you’d be placing limits on the freedom to modify as well.

> Instead, the data information and code requirements allow Open Source AI systems to be forked by third-party AI builders downstream using the same information as the original developers.

This is internally inconsistent: the “same information as the original developers” IS the data, not metadata (data about the data aka “data information”).

@Shamar [nailed it](https://discuss.opensource.org/t/the-osaid-requires-training-data-to-be-shared/619/5) in that “as long as the training and testing data are available to the public under the same terms that allowed their usage from the builders in the first place, we can still count the AI system as “Open Source AI”, since the 4 freedoms are still granted even if the builders cannot directly distribute them.”

> These forks could include removing non-public or non-open data from the training dataset, in order to retrain a new Open Source AI system on fully public or open data.

You cannot remove non-public or non-open data from the training dataset — which would be a very common form of exercise of the freedom to modify — if you don’t have the training dataset.

Per @thesteve0 in [Model Weights is not enough for Open Source AI](https://blog.techravenconsulting.com/model-weights-is-not-enough-for-open-source-ai/), prove it with a “demonstration showing that both techniques produce the same model weights.”

 ![a-diagram-of-a-model-fitting-description-automati (1)](https://canada1.discourse-cdn.com/flex011/uploads/opensource2/original/1X/345c22855ddfbba319bb2b1cd83bde8bfad6d42f.png)

> [@stefano](#):
>
> This insight was shaped by input from the community and experts who joined our study groups and voted on various approaches.

The vote was [misinterpreted](https://discuss.opensource.org/t/we-heard-you-lets-focus-on-substantive-discussion/589/25) and [demands the data](https://discuss.opensource.org/t/we-heard-you-lets-focus-on-substantive-discussion/589/9?). So does the straw poll I’m running asking “What is the “preferred form” in which a practitioner would modify a model?” for which 100% of the votes are going to _Data_ rather than _Model_.

Perhaps this was pre-posted before the developments of the last days, or was intended to reflect the state of affairs as at the [Open Source Summit Europe](https://events.linuxfoundation.org/archive/2024/open-source-summit-europe/) on 16-18 September, but if not it feels like we’re [going forward internally](https://discuss.opensource.org/t/the-osaid-requires-training-data-to-be-shared/619) only to go [backwards externally](https://opensource.org/blog/a-journey-toward-defining-open-source-ai-presentation-at-open-source-summit-europe).

---

_[View the full topic](https://discuss.opensource.org/t/a-journey-toward-defining-open-source-ai-presentation-at-open-source-summit-europe/624)._
