# The Missing Third Leg: Training Data Excluded from Open Source AI Definition by \[Co-\]Design

**URL:** <https://discuss.opensource.org/t/the-missing-third-leg-training-data-excluded-from-open-source-ai-definition-by-co-design/611>\
**Category:** Open Source AI\
**Created:** [September 28, 2024, 9:54pm UTC](https://discuss.opensource.org/t/the-missing-third-leg-training-data-excluded-from-open-source-ai-definition-by-co-design/611 "2024-09-28T21:54:15Z")\
**Posts on this page:** 1\
**Showing post:** 8

<div class="post-metadata">

**Author:** ![anon18632855](https://avatars.discourse-cdn.com/v4/letter/a/dfb087/32.png) [@anon18632855](https://discuss.opensource.org/u/anon18632855)\
**Post date:** [September 30, 2024, 8:45am UTC](https://discuss.opensource.org/t/the-missing-third-leg-training-data-excluded-from-open-source-ai-definition-by-co-design/611/8 "2024-09-30T08:45:08Z")

</div>

> [@quaid](#):
>
> Completely setting aside nefarious actors, it seems that having the dual-branding (like D+/D-) gives a way for model creators from research, academia, NGOs, and startups to participate in the OSAI ecosystem on a more equal footing, yes?

Yes. Per my [reply](https://discuss.opensource.org/t/is-lgpl-really-a-precedent-for-an-open-washing-ai-definition/617/2) to @Shamar, the precedent is for the licence that is limited for pragmatic reasons (e.g., LGPL) to adopt the alternative branding while still protecting the four freedoms. Alternatively, any compromise could be temporary under a single brand, but it may be difficult or impossible to revoke later.

“The proposal is that we acknowledge that taking a purist approach to data (e.g., demanding open data licenses) will drastically limit the number of candidates for certification”, to your points above, rather “requiring data (analogous to the proprietary code [under the LGPL precedent]) be accessible when building (training) the Open Source licensed redistributable software (model), in order to protect the four freedoms”.

> [@lumin](#):
>
> Once the company behind it goes down, nobody else in the world would be able to reproduce this work in order to continue maintaining the work or the piece of AI software.

Freeware (free as in beer, not as in freedom) is what we used to call software that was made available in binary-only form to _use_ and _share_ but not _study_ or _modify_, as you know, and this is one of the main reasons people are reluctant (or prohibited by policies) to rely on it. The [current draft](https://opensource.org/deepdive/drafts/open-source-ai-definition-draft-v-0-0-9) is basically freeware for AI, and this impinges on the freedom to even _use_ the software.

> [@lumin](#):
>
> This is not what a free software community looks like.

No, it’s not, which is why I consider this a train worth throwing ourselves in front of.

---

_[View the full topic](https://discuss.opensource.org/t/the-missing-third-leg-training-data-excluded-from-open-source-ai-definition-by-co-design/611)._
