# Proposal to handle Data Openness in the Open Source AI definition \[RFC\]

**URL:** <https://discuss.opensource.org/t/proposal-to-handle-data-openness-in-the-open-source-ai-definition-rfc/561>\
**Category:** Open Source AI\
**Tags:** draft\
**Created:** [September 12, 2024, 1:19am UTC](https://discuss.opensource.org/t/proposal-to-handle-data-openness-in-the-open-source-ai-definition-rfc/561 "2024-09-12T01:19:17Z")\
**Posts on this page:** 2\
**Page:** 2

<div class="post-metadata">

**Author:** ![anon18632855](https://avatars.discourse-cdn.com/v4/letter/a/dfb087/32.png) [@anon18632855](https://discuss.opensource.org/u/anon18632855)\
**Post date:** [October 6, 2024, 6:46pm UTC](https://discuss.opensource.org/t/proposal-to-handle-data-openness-in-the-open-source-ai-definition-rfc/561/21 "2024-10-06T18:46:01Z")

</div>

> [@quaid](#):
>
> And now [Open has a definition](https://opendefinition.org) based on time-tested concepts, and it has proven useful.

TIL “Open” is now well-defined covering 3 of 4 freedoms (_study_ being absent):

> “Open data and content can be **freely used, modified, and shared** by **anyone** for **any purpose** ”

And RC1 doesn’t meet even this looser standard. In other words, the Open Source AI Definition (OSAID) is not, in fact, _Open_ according to its accepted definition.

Or at least those claiming it does have not shown it does, while those who dispute the claim have shown counterexamples.

> [@quaid](#):
>
> I no longer think it’s possible or desirable to have a double-designation (“secondary brand”) for an Open-related definition if _the secondary designation is antithetical to the letter and spirit of Open_.

Agreed. Either you fully protect the four _essential_ freedoms, or you do not.

While _study_ does start to venture into the Open Science realm, it was deemed important enough to be considered _essential_.

> [@quaid](#):
>
> They have plenty of other Opens they can lay some claim to, let that be enough.

Agreed. Adding a definition that is both difficult to meet (in the spirit of the rules) and yet easy to circumvent (to the letter of rules) does not help and likely hurts our cause.

---

<div class="post-metadata">

**Author:** ![gvlx](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/gvlx/32/381_2.png) [@gvlx](https://discuss.opensource.org/u/gvlx)\
**Post date:** [October 8, 2024, 4:54pm UTC](https://discuss.opensource.org/t/proposal-to-handle-data-openness-in-the-open-source-ai-definition-rfc/561/22 "2024-10-08T16:54:22Z")

</div>

Adding one more document into the discussion.

[1]‘Model AI Governance Framework for Generative AI’, AI Verify Foundation. Accessed: Oct. 08, 2024. [Online]. Available: [MGF for GenAI – AI Verify Foundation](https://aiverifyfoundation.sg/resources/mgf-gen-ai/)

[2] ‘Singapore proposes framework to foster trusted Generative AI development’, Infocomm Media Development Authority. Accessed: Oct. 08, 2024. [Online]. Available: [Model AI Governance Framework 2024 - Press Release | IMDA](https://www.imda.gov.sg/resources/press-releases-factsheets-and-speeches/press-releases/2024/public-consult-model-ai-governance-framework-genai)

[1, p.11]

> ## Facilitating Access to Quality Data
> 
> As an overall hygiene measure at an organisational level, it would be good discipline for AI developers to undertake data quality control measures and adopt general best practices in data governance, including annotating training datasets consistently and accurately, and using data analysis tools to facilitate data cleaning (e.g., debiasing and removing inappropriate content).
> 
> Globally, it is worth considering a concerted effort to expand the available pool of trusted datasets. Reference datasets are important tools in both AI model development (e.g., for fine-tuning) as well as benchmarking and evaluation.21 Governments can also consider working with their local communities to curate a repository of representative training datasets for their specific context (e.g., in low resource languages). This helps to improve the availability of quality datasets that reflect the cultural and social diversity of a country, which in turn supports the development of safer and more culturally representative models.

[Previous page](https://discuss.opensource.org/t/proposal-to-handle-data-openness-in-the-open-source-ai-definition-rfc/561.md?page=1)
