# Definition Validation: Seeking Volunteers

**URL:** <https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320>\
**Category:** Open Source AI\
**Tags:** process\
**Created:** [May 1, 2024, 5:37pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320 "2024-05-01T17:37:29Z")\
**Posts on this page:** 10\
**Page:** 1

<div class="post-metadata">

**Author:** ![Mer](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/mer/32/102_2.png) [@Mer](https://discuss.opensource.org/u/Mer)\
**Post date:** [May 1, 2024, 5:37pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/1 "2024-05-01T17:37:29Z")

</div>

**CONTEXT** Last week Stefano [announced](https://discuss.opensource.org/t/draft-v-0-0-8-of-the-open-source-ai-definition-is-available-for-comments/315) [v. 0.0.8](https://hackmd.io/@opensourceinitiative/osaid-0-0-8), the first feature complete version of the Open Source AI Definition (OSAID).

**ASK** We are now seeking volunteers to validate that definition by using it to review additional AI systems that are self-described as open. If you are interested, please comment below or DM me.

**TIME** We would like to complete this validation task by **Monday, May 20th**.

**TASK** You will have a [spreadsheet](https://docs.google.com/spreadsheets/d/1SCT2K-8xvTIbozpoXGuvf1NVcB54vY8p2fKkJlyk4KM/edit#gid=1544910696) (example below) in which you locate and link to the license, research paper, or other document that grants rights or provides information for each required component. You will then indicate in each cell whether the document Allows or Restricts the ability to study use, modify, or share that component.

 ![Screenshot 2024-05-01 at 10.08.59 AM](https://canada1.discourse-cdn.com/flex011/uploads/opensource2/original/1X/f83740c79073e5b4035bfc990ad6685886e62188.jpeg)  
Example: spreadsheet used to review Llama 2 according OSAID v. 0.0.6

**SYSTEMS** We are interested in reviewing about ten self-described open AI systems as part of this definitional process. Four (marked \*) have already been reviewed by the [workgroups](https://discuss.opensource.org/t/report-on-working-group-recommendations/247). Below are some more systems we are interested in being part of this validation task. If there is another system you would like to review, please say so.

1. **Arctic** Jesús M. Gonzalez-Barahona
2. **BLOOM** \* Danish Contractor, Jaan Li
3. **Falcon** Casey Valk, Jean-Pierre Lorre
4. **Grok** Victor Lu, Karsten Wade
5. **Llama 2** \* Davide Testuggine, Jonathan Torres, Stefano Zacchiroli, Victor Lu
6. **Mistral** Mark Collier, Jean-Pierre Lorre, Cailean Osborne
7. **OLMo** Amanda Casari, Abdoulaye Diack
8. **OpenCV** \* Rasim Sen
9. **Phi-2** Seo-Young Isabelle Hwang
10. **Pythia** \* Seo-Young Isabelle Hwang, Stella Biderman, Hailey Schoelkopf, Aviya Skowron
11. **T5** Jaan Li

**TO VOLUNTEER** Comment below or DM me if you would like to volunteer. Anyone who can complete the task may do so, either solo or as a group. If you are a creator or advisor on the system you are reviewing, we will also need to identity an unaffiliated individual to review the system, so please disclose that when you volunteer. Your name and organizational affiliation will also be made public as part of our transparency policy.

Women, trans, and non-binary folx, black, indigenous, latine/o/a, and other people of color, immigrants, people with disabilities, and people from poor and working class backgrounds are encouraged to respond.

Thanks 🙂

EDIT: I’ll update the list above with volunteer names so it’s clear where there is greatest need.  
(last update: May 14th @ 11:03 am PDT)

---

<div class="post-metadata">

**Author:** ![amcasari](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/amcasari/32/247_2.png) [@amcasari](https://discuss.opensource.org/u/amcasari)\
**Post date:** [May 2, 2024, 1:30pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/2 "2024-05-02T13:30:21Z")

</div>

I volunteer to help with a review!

I have a conflict of interest with #10, T5 (same employer), but could take the lead or assist on any others.

I would prefer to start w/ OLMo, if that isn’t taken yet.

---

<div class="post-metadata">

**Author:** ![Mer](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/mer/32/102_2.png) [@Mer](https://discuss.opensource.org/u/Mer)\
**Post date:** [May 2, 2024, 4:29pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/3 "2024-05-02T16:29:45Z")

</div>

Thank you, @amcasari! You’ve got review on OLMo. I’m going to make up the new review spreadsheet today and will email you with further details.

---

<div class="post-metadata">

**Author:** ![stefano](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/stefano/32/4_2.png) [@stefano](https://discuss.opensource.org/u/stefano)\
**Post date:** [May 3, 2024, 12:06pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/4 "2024-05-03T12:06:00Z")

</div>

Here is another one claiming to be “truly open”: Can we get someone to review it, please?

> **[Snowflake Arctic - LLM for Enterprise AI](https://www.snowflake.com/blog/arctic-open-efficient-foundation-language-models-snowflake/)**
>
> Snowflake Arctic from the Snowflake AI research team is a top-tier enterprise LLM that pushes the frontiers of cost-effective training and openness.

---

<div class="post-metadata">

**Author:** ![Aspie96](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/aspie96/32/73_2.png) [@Aspie96](https://discuss.opensource.org/u/Aspie96)\
**Post date:** [May 4, 2024, 2:33am UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/5 "2024-05-04T02:33:58Z")

</div>

> [@Mer](#):
>
> If there is another system you would like to review, please say so.

There are two kinds of systems I don’t see listed that I would like to see.

1. Non-foundational models.
2. Systems which are not based on deep learning (and, preferably, not even on machine learning).

That said, if I can only suggest specific systems, I have at least two:

- [Open Image Denoise](https://www.openimagedenoise.org/), by Intel.
- NLP models such as the [Stanford Log-linear Part-Of-Speech Tagger](https://nlp.stanford.edu/software/tagger.html).

---

<div class="post-metadata">

**Author:** ![stefano](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/stefano/32/4_2.png) [@stefano](https://discuss.opensource.org/u/stefano)\
**Post date:** [May 6, 2024, 9:25am UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/6 "2024-05-06T09:25:44Z")

</div>

> [@Aspie96](#):
>
> That said, if I can only suggest specific systems, I have at least two:

Good thinking, thanks! I think you can get started by cloning [this table structure](https://docs.google.com/spreadsheets/d/1SCT2K-8xvTIbozpoXGuvf1NVcB54vY8p2fKkJlyk4KM/edit?usp=sharing) and fill in the details.

---

<div class="post-metadata">

**Author:** ![quaid](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/quaid/32/277_2.png) [@quaid](https://discuss.opensource.org/u/quaid)\
**Post date:** [May 13, 2024, 6:18pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/7 "2024-05-13T18:18:47Z")

</div>

Heya Mer,

I love this exercise! I’ve been doing this lightly across systems recently but not to this same depth and level of certainty. In fact, Snowflake Arctic is one I looked more closely at, so I’m happy to help with this one if needed.

Otherwise, I have no preferences and will help wherever needed. Would you like to choose several for me, prioritize them if possible? I’m not aware of my having any affiliation with these systems, and Open Community Architects (OCA) is a neutral consultancy in those regards (aside from a bias for Open.)

- Karsten (quaid)

---

<div class="post-metadata">

**Author:** ![Mer](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/mer/32/102_2.png) [@Mer](https://discuss.opensource.org/u/Mer)\
**Post date:** [May 13, 2024, 6:36pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/8 "2024-05-13T18:36:09Z")

</div>

Thanks, @quaid. I think Arctic is well-covered, but let me see if anyone else needs help. If not, you can just pick a system that works for you. I’ll get back to you soon.

---

<div class="post-metadata">

**Author:** ![Mer](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/mer/32/102_2.png) [@Mer](https://discuss.opensource.org/u/Mer)\
**Post date:** [May 14, 2024, 6:02pm UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/9 "2024-05-14T18:02:45Z")

</div>

@quaid I got no requests from the current volunteers, so I’m going to put you on Grok, which is a new system with only 1 volunteer. Please DM me your email address so I can add you to the reviewer chat and spreadsheet.

---

<div class="post-metadata">

**Author:** ![Aspie96](https://yyz1.discourse-cdn.com/flex011/user_avatar/discuss.opensource.org/aspie96/32/73_2.png) [@Aspie96](https://discuss.opensource.org/u/Aspie96)\
**Post date:** [May 16, 2024, 4:53am UTC](https://discuss.opensource.org/t/definition-validation-seeking-volunteers/320/10 "2024-05-16T04:53:22Z")

</div>

Two models I hadn’t thought about mentioning previously, but are rather interesting, are Segment Anything, by Meta and Whisper, by OpenAI.

What makes them interesting is the fact that they sparked a lot of attention when they were published and they are built by companies that usually make in-house models (OpenAI) or often release models under proprietary licenses (Meta), but in this case were released under open source licenses (both code and weigths): Apache 2.0 and MIT license respectively.
