# Draft v.0.0.7 of the Open Source AI Definition is available for comments

**URL:** <https://discuss.opensource.org/t/draft-v-0-0-7-of-the-open-source-ai-definition-is-available-for-comments/298>\
**Category:** Open Source AI\
**Tags:** draft\
**Created:** [April 12, 2024, 4:10pm UTC](https://discuss.opensource.org/t/draft-v-0-0-7-of-the-open-source-ai-definition-is-available-for-comments/298 "2024-04-12T16:10:36Z")\
**Posts on this page:** 1\
**Showing post:** 8

<div class="post-metadata">

**Author:** ![zack](https://avatars.discourse-cdn.com/v4/letter/z/f6c823/32.png) [@zack](https://discuss.opensource.org/u/zack)\
**Post date:** [April 17, 2024, 8:13am UTC](https://discuss.opensource.org/t/draft-v-0-0-7-of-the-open-source-ai-definition-is-available-for-comments/298/8 "2024-04-17T08:13:44Z")

</div>

> [@stefano](#):
>
> The issue is not that “most models will not meet the definition” but that _none_ will do.

I’m not sure if you meant that in absolute terms, but for what is worth, this is not true.

For instance, in the domain of LLMs for Code, [StarCoder2](https://github.com/bigcode-project/starcoder2) is a state-of-the-art model whose training dataset is redistributable (and redistributed).

I’m less familiar with other application domains, but ML models trained _only_ on data sources like Wikipedia, Wikidata, Wikimedia Commons, etc., could also easily redistribute their training datasets.

---

_[View the full topic](https://discuss.opensource.org/t/draft-v-0-0-7-of-the-open-source-ai-definition-is-available-for-comments/298)._
