# Welcome diverse approaches to training data within a unified Open Source AI Definition

**URL:** <https://discuss.opensource.org/t/welcome-diverse-approaches-to-training-data-within-a-unified-open-source-ai-definition/531>\
**Category:** Open Source AI\
**Created:** [August 30, 2024, 5:24pm UTC](https://discuss.opensource.org/t/welcome-diverse-approaches-to-training-data-within-a-unified-open-source-ai-definition/531 "2024-08-30T17:24:57Z")\
**Posts on this page:** 1\
**Showing post:** 14

<div class="post-metadata">

**Author:** ![Shamar](https://avatars.discourse-cdn.com/v4/letter/s/94ad74/32.png) [@Shamar](https://discuss.opensource.org/u/Shamar)\
**Post date:** [September 15, 2024, 9:44am UTC](https://discuss.opensource.org/t/welcome-diverse-approaches-to-training-data-within-a-unified-open-source-ai-definition/531/14 "2024-09-15T09:44:34Z")

</div>

> [@stefano](#):
>
> - if the system is trained on [unshareable non-public training data](https://opensource.org/blog/community-input-drives-the-new-draft-of-the-open-source-ai-definition), allow the downstream developers to use open data to fine tune that model
> 
> In update #2 you seem to be arguing that training data is to trained model weights as software source code is to binary code.

While I’d support the first two change sugestions, and I totally agree that **training data is to trained model weights as software source code is to binary executable** , I can’t see how a “system trained on [unshareable non-public training data](https://opensource.org/blog/community-input-drives-the-new-draft-of-the-open-source-ai-definition)” could match an “Open **Source** AI” definition.

Such system would not provide the freedom to study the system and would limit the freedom to modify the model in a huge way.

Open Source grants the **freedom to study and modify** or just the **freedom to fine tune**?

---

_[View the full topic](https://discuss.opensource.org/t/welcome-diverse-approaches-to-training-data-within-a-unified-open-source-ai-definition/531)._
