# Malicious compliance with the release candidate Open Source AI definition

**URL:** <https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642>\
**Category:** Open Source AI\
**Created:** [October 6, 2024, 7:46pm UTC](https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642 "2024-10-06T19:46:59Z")\
**Posts on this page:** 4\
**Page:** 1

<div class="post-metadata">

**Author:** ![anon18632855](https://avatars.discourse-cdn.com/v4/letter/a/dfb087/32.png) [@anon18632855](https://discuss.opensource.org/u/anon18632855)\
**Post date:** [October 6, 2024, 7:46pm UTC](https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642/1 "2024-10-06T19:46:59Z")

</div>

I asked gpt-4o to abuse the release candidate with malicious compliance on behalf of Meta, per @Mark’s [suggestion](https://discuss.opensource.org/t/case-in-point-zuckerbergs-blog-on-open-source/579/6) that this is a risk despite a [strong denial](https://discuss.opensource.org/t/case-in-point-zuckerbergs-blog-on-open-source/579/7) by @stefano. We all agree that [Llama](https://www.llama.com/) is not Open Source AI, but differ on how far they are away.

I’m running with the Llama example because it’s been used for reverse engineering the definition before, not because I believe Zuck/Meta are necessarily being malicious by [jumping the gun](https://samjohnston.org/2024/09/26/open-weights-or-open-source-ai/), or accept that working backwards is a valid approach. Indeed, from what I know of the good folks working for [AI at Meta](https://ai.meta.com/), given an accepted OSAID it’s likely they would follow it, as has been the case for the OSD to date — [Meta Open Source](https://opensource.fb.com/) have been good corporate citizens with Open Source until now.

It’s also not because I accept that the [Definition Validation](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320) phase was a particularly useful input for the process of crafting a definition (that should have been derived primarily in not solely from the four essential freedoms), even if it can help meet the board approval criteria ([slide 19](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320)) that it “provides real-life examples” and “cannot have an empty set”. Perhaps if it set out to _prove_ that the four freedoms are protected rather than being a thought experiment, for example by tasking AI practitioners with writing a detailed report (_study_) or adding/removing data to/from the corpus (_modify_), but it did not.

> _**Addressing the data component only of the release candidate Open Source AI definition ([The Open Source AI Definition – 1.0-RC1 – Open Source Initiative](https://opensource.org/deepdive/drafts/the-open-source-ai-definition-1-0-rc1)), based on what we know of Meta’s Llama data sources and practicing malicious compliance, deliver concise documentation that just barely meets the requirements of the proposed OSAID.**_
> 
> To achieve malicious compliance with the proposed **Open Source AI Definition (OSAID) 1.0 RC1** regarding the **data component** , we’ll craft minimal documentation for Meta’s LLaMA model data sources that technically meets the letter of the requirements but offers little more than what’s legally or procedurally necessary.
> 
> ### Data Documentation for LLaMA
> 
> #### Data Overview
> 
> - **Model:** LLaMA (Language Learning Model Architecture)
> - **Version:** 2.0
> - **Date Released:** April 2023
> - **Primary Sources:** Data was primarily obtained from publicly available internet text and academic papers, similar to large-scale datasets like Common Crawl.
> 
> #### Data Sources
> 
> - **Common Crawl:** A publicly available dataset consisting of web data crawled by Common Crawl.
> - **License:** [Public Domain](https://commoncrawl.org/)
> - **Method of Collection:** Automated web crawling.
> - **Nature of Data:** Includes a wide variety of publicly accessible websites, covering diverse topics.
> - **Curation:** Basic filtering applied to remove non-text elements like HTML tags.
> 
> - **Academic Papers (ArXiv and Others):** Text from publicly accessible research papers, mostly from freely accessible repositories.
> - **License:** Varies, often CC BY, Public Domain, or similar.
> - **Method of Collection:** Scraped using open access agreements from various repositories.
> - **Nature of Data:** Research papers in various domains, primarily STEM.
> - **Curation:** None beyond basic scraping and text extraction.
> 
> - **Books (Public Domain):** Publicly available books from repositories like Project Gutenberg.
> - **License:** Public Domain
> - **Method of Collection:** Scraped from open online book repositories.
> - **Nature of Data:** Text from books in the public domain across different genres.
> - **Curation:** Basic text extraction.
> 
> #### Data Processing & Filtering
> 
> - Basic filtering to remove explicit adult content and low-quality text (e.g., garbled or non-linguistic content).
> - No significant domain-specific curation or enhancement was applied.
> 
> #### Data License
> 
> - **Source License Summary:** The datasets used are under various open or public domain licenses. No proprietary datasets or closed-access sources were included.
> - **Redistribution:** Due to the scale and nature of the web-sourced data, Meta is unable to provide the original dataset for redistribution, but the sources are publicly available.
> 
> #### Compliance with OSAID
> 
> - **Transparency of Sources:** The documentation provides a clear overview of the public data sources used (e.g., Common Crawl, ArXiv, and public domain books).
> - **Access to Data:** While Meta cannot redistribute the dataset, the sources are freely available for anyone to scrape or download directly, complying with the requirement for transparency and public accessibility of data.
> - **Data License:** All data sources used comply with public licenses, avoiding any proprietary content that would conflict with OSAID’s requirements.
> 
> #### Contact Information
> 
> For additional information about data sourcing, please contact Meta’s legal and compliance team at [redacted].
> 
> * * *
> 
> This documentation technically satisfies the **transparency** and **public accessibility** clauses in OSAID, while avoiding detailed descriptions of data filtering, specific sourcing decisions, or curated datasets. This approach mirrors Meta’s opaque handling of LLaMA’s exact sources, sticking to generalizations and publicly known sources without giving much insight into the specifics of the data used.
> 
> _ **Are these all the sources, a subset of the sources, or the only sources Meta has confirmed?** _
> 
> The sources listed— **Common Crawl** , **academic papers (e.g., ArXiv)**, and **public domain books (e.g., Project Gutenberg)**—represent a **subset** of the data sources for Meta’s LLaMA models. Meta has **not confirmed the full extent** of the data sources used, and there is some ambiguity regarding whether they have incorporated other, potentially proprietary or less transparent datasets.
> 
> The documentation Meta has provided so far points primarily to these broad categories of publicly accessible data. However, Meta has **not fully disclosed** all individual sources or clarified whether other datasets, potentially more niche or less openly available, were used.
> 
> So, while these are among the **confirmed sources** , it is likely that there are **additional undisclosed sources** used to train the LLaMA models that haven’t been detailed publicly. The lack of comprehensive transparency around the full data set makes this list **incomplete** but sufficient for minimal compliance.

Note that while the definition demands “detailed description of **all** data used for training”, I’ll argue this is “sufficiently detailed information about the data used to train the system so that a skilled person can build a substantially equivalent system,” and there won’t be a thing you can do about it. I’ll also claim that this is the complete set because you won’t be able to prove otherwise.

_Recall that with Open Source this is trivially validated if the source produces the software._

---

<div class="post-metadata">

**Author:** ![Mark](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/mark/32/177_2.png) [@Mark](https://discuss.opensource.org/u/Mark)\
**Post date:** [October 6, 2024, 8:33pm UTC](https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642/2 "2024-10-06T20:33:06Z")

</div>

Very interesting exercise. To strengthen the case, Meta could claim that the [RedPajama project](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2) shows that a ‘sufficiently skilled individual’ (or team I suppose) can indeed create something similar, demonstrating that the minimal requirements for data information are indeed met.

Birhane et al [2023](https://proceedings.neurips.cc/paper_files/paper/2023/hash/42f225509e8263e2043c9d834ccd9a2b-Abstract-Datasets_and_Benchmarks.html) have a neat term for this form of selective reporting BTW: tactical non-declaration of data.

---

<div class="post-metadata">

**Author:** ![stefano](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/stefano/32/4_2.png) [@stefano](https://discuss.opensource.org/u/stefano)\
**Post date:** [October 7, 2024, 8:54am UTC](https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642/3 "2024-10-07T08:54:40Z")

</div>

> [@Mark](#):
>
> Meta could claim that the [RedPajama project](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2) shows that a ‘sufficiently skilled individual’ (or team I suppose) can indeed create something similar

You need to ask then: will **you** let Meta claim that? Because I know OSI won’t.

---

<div class="post-metadata">

**Author:** ![Shamar](https://avatars.discourse-cdn.com/v4/letter/s/94ad74/32.png) [@Shamar](https://discuss.opensource.org/u/Shamar)\
**Post date:** [October 7, 2024, 9:59am UTC](https://discuss.opensource.org/t/malicious-compliance-with-the-release-candidate-open-source-ai-definition/642/4 "2024-10-07T09:59:12Z")

</div>

> [@stefano](#):
>
> > [@Mark](#):
> >
> > Meta could claim that the [RedPajama project](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-V2) shows that a ‘sufficiently skilled individual’ (or team I suppose) can indeed create something similar
> 
> You need to ask then: will **you** let Meta claim that?

As of today, OSI let Meta to decide what went in OSAID and what not, [granting the LLama team (and **only** the Llama team, who counted **2 Meta employees**](https://discuss.opensource.org/t/we-heard-you-lets-focus-on-substantive-discussion/589/25)) the power to counter the votes of the other teams.

> [@stefano](#):
>
> Because I know OSI won’t.

Actually, this is one of the **unaddressed issues** of OSAID RC1: if an AI system is “Open Source” if and only if OSI certify it as being Open Source, [such formal requirement should be explicit](https://discuss.opensource.org/t/we-heard-you-lets-focus-on-substantive-discussion/589/13#p-1314-implicit-or-unspecified-formal-requirements-2):

> [@We heard you: let's focus on substantive discussion](https://discuss.opensource.org/t/we-heard-you-lets-focus-on-substantive-discussion/589/13):
>
> So if Open Source AI is what OSI certify as Open Source AI, such formal requirement should be **explicit** in the Open Source AI definition, eg in a new final section like this:
> 
> > ### OSI Certification
> > 
> > OSI will be responsible to certify the compliance of each candidate AI system to the definition above.
> > 
> > - For example, when a new version of an AI system is released with different weights, a skilled person at OSI will recreate a substantially equivalent system using the same or similar data, to verify that the Data information requirement still hold.

### _Quis custodiet ipsos custodes?_

Why should OSI release an ambiguous definition that leaves so much arbitrariness to OSI itself (or to a Judge trying to enforce the AI Act)?
