# Open Source AI needs to require data to be viable

**URL:** <https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351>\
**Category:** Open Source AI\
**Created:** [May 19, 2024, 6:03pm UTC](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351 "2024-05-19T18:03:48Z")\
**Posts on this page:** 1\
**Showing post:** 5

<div class="post-metadata">

**Author:** ![Danish\_Contractor](https://avatars.discourse-cdn.com/v4/letter/d/c57346/32.png) [@Danish\_Contractor](https://discuss.opensource.org/u/Danish_Contractor)\
**Post date:** [May 27, 2024, 2:00pm UTC](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351/5 "2024-05-27T14:00:07Z")

</div>

No I understand that – I’m trying to see if we could incentivize more sharing via the Open Source AI draft; currently this draft effectively just says – remove usage restrictions from an OpenRAIL-MS license and we have Open Source AI.

And lets say we do that – if I wanted to modify a given “Open Source AI” system as defined by the draft, in a way that excludes (for example) all wikipedia content (assuming that was part of training data); I cant make that modification.  
Is that an acceptable constraint on Openness in Open Source AI ?

There’s nothing wrong with open sourcing a _component_ of an AI system (weights in this case) as opposed to the full system but then maybe we should make that clearer in the language.

I had made a suggestion here in this thead: [Recognising Open Source "Components" of an AI System](https://discuss.opensource.org/t/recognising-open-source-components-of-an-ai-system/207)

---

_[View the full topic](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351)._
