# On the current definition of Open Source AI and the state of the data commons

**URL:** <https://discuss.opensource.org/t/on-the-current-definition-of-open-source-ai-and-the-state-of-the-data-commons/559>\
**Category:** Open Source AI\
**Created:** [September 11, 2024, 4:59pm UTC](https://discuss.opensource.org/t/on-the-current-definition-of-open-source-ai-and-the-state-of-the-data-commons/559 "2024-09-11T16:59:20Z")\
**Posts on this page:** 1\
**Showing post:** 15

<div class="post-metadata">

**Author:** ![Shamar](https://avatars.discourse-cdn.com/v4/letter/s/94ad74/32.png) [@Shamar](https://discuss.opensource.org/u/Shamar)\
**Post date:** [September 15, 2024, 1:32pm UTC](https://discuss.opensource.org/t/on-the-current-definition-of-open-source-ai-and-the-state-of-the-data-commons/559/15 "2024-09-15T13:32:14Z")

</div>

Sorry, but I can’t follow your argument.

If the data sources are available and legally usable to train the LLM, a simple link with versioning and a sha512sum would be enough.

We are exactly in the same scenario [you proposed above](https://discuss.opensource.org/t/on-the-current-definition-of-open-source-ai-and-the-state-of-the-data-commons/559/8) and the same solution apply.

> [@Shamar](#):
>
> if the developers provide enough information about how to retrieve and process the data “so that a skilled person can recreate an **exact copy** of the system using the same data”, they can legitimately call their system “Open Source AI”.

The point remains: it’s perfectly possible to create a LLM that comply to a Open Source AI definition that requires the availability of the data used to train the model’s weights, "so that a skilled person can recreate an exact copy of the system using the same data”.

> [@shujisado](#):
>
> This is the last reply to this thread. The thread is a bit long.

Well, it’s not the [longest thread we have joined so far](https://discuss.opensource.org/t/draft-v-0-0-9-of-the-open-source-ai-definition-is-available-for-comments/513/17), but it has been a deep exchange on the supposed limits of a proper Open Source AI definition, that proved its applicability to interesting corner cases.

Anyway, thanks for the conversation and if you have further doubts or suggestions, I’d be happy to discusss them!

---

_[View the full topic](https://discuss.opensource.org/t/on-the-current-definition-of-open-source-ai-and-the-state-of-the-data-commons/559)._
