# Open Source AI needs to require data to be viable

**URL:** <https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351>\
**Category:** Open Source AI\
**Created:** [May 19, 2024, 6:03pm UTC](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351 "2024-05-19T18:03:48Z")\
**Posts on this page:** 1\
**Showing post:** 34

<div class="post-metadata">

**Author:** ![stellaathena](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/stellaathena/32/185_2.png) [@stellaathena](https://discuss.opensource.org/u/stellaathena)\
**Post date:** [June 6, 2024, 9:18pm UTC](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351/34 "2024-06-06T21:18:36Z")

</div>

> [@juliaferraioli](#):
>
> Great call out @nick on the license change for [dolma](https://huggingface.co/datasets/allenai/dolma). Given that, I do think that OLMo would qualify as open source based on the criteria that I posit is necessary. Pythia still would not as the data are not licensed.

Dolma contains data that is not being used in a fashion consistent with its license, just like the Pile does. They didn’t go through C4 and validate the licensing of everything in it.

---

_[View the full topic](https://discuss.opensource.org/t/open-source-ai-needs-to-require-data-to-be-viable/351)._
