# Draft v.0.0.9 of the Open Source AI Definition is available for comments

**URL:** <https://discuss.opensource.org/t/draft-v-0-0-9-of-the-open-source-ai-definition-is-available-for-comments/513>\
**Category:** Open Source AI\
**Tags:** draft\
**Created:** [August 22, 2024, 12:30pm UTC](https://discuss.opensource.org/t/draft-v-0-0-9-of-the-open-source-ai-definition-is-available-for-comments/513 "2024-08-22T12:30:06Z")\
**Posts on this page:** 1\
**Showing post:** 30

<div class="post-metadata">

**Author:** ![Shamar](https://avatars.discourse-cdn.com/v4/letter/s/94ad74/32.png) [@Shamar](https://discuss.opensource.org/u/Shamar)\
**Post date:** [September 23, 2024, 11:31pm UTC](https://discuss.opensource.org/t/draft-v-0-0-9-of-the-open-source-ai-definition-is-available-for-comments/513/30 "2024-09-23T23:31:31Z")

</div>

Hi @mjbommar, @samj and all, I’m sorry for the delay but I’ve just got back access to my account [after being silenced for the proposal to separate concerns between source data and processing information](https://discuss.opensource.org/t/rfc-separating-concerns-between-source-data-and-processing-information/568/3).

The idea was simply to require training data to **be** available under the same terms that allowed their use in training in the first place whenever we cannot require them to be **made** available, so that no legal issue arise from the requirement to distribute them.

But probably my English is way worse than I suspect, so the thread is closed without any comment, but for a few questions I’m not allowed to answer.

> [@mjbommar](#):
>
> It’s not well-accepted that you can deterministically recreate “realistic” scaled models.  
> […]  
> this is all ignoring corruption that occurs.

> [@Training data access](https://discuss.opensource.org/t/training-data-access/152/53):
>
> That level of reproducibility (byte-to-byte) is extremely difficult with CUDA, even with the identical random number generators. Different generations of GPUs will give you different results due to the hardware float point implementation and the instruction sets.

Race conditions are bugs, not features.  
So are data corruptions in RAM.

While they might constitute a sustainable technical dept in some situations, they shouldn’t be “normalized”. In fact, a lot of work as been done to achieve determinism since the GTC 2019 talk [_Determinism in Deep Learning_](http://bit.ly/determinism-in-deep-learning).

I’d argue that it’s **always** possible (accepting a performance toll) to get exact training reproducibility (on the same hardware) with proper design.

But let’s suppose you face some technical limitation that affects all the existing hardware and that something along the line “sort floats before adding them” cannot fix.

The source of randomness can be identified and recorded, so that people can still **study** exactly how the training process produced the weights from the original data. At worst, you’d have to record the whole training process. Heavy, slow, expensive (without the proper environment) but not impossible.

Now the point is: why should you?

> [@samj](#):
>
> As [discussed](https://discuss.opensource.org/t/training-data-access/152/53), byte-for-byte reproducibility is neither an achievable nor necessarily useful goal.

In several AI systems, you don’t need to do much to achieve reproducibility (the training process is inherently reproducible), but in a few ones with huge social impact, it’s needed to avoid the **open washing** of **black boxes**.

Sure, as @samj pointed out, the open source definition does not mandate reproducible builds and it’s normal to get different binaries when compiling a project with different flags.

But such non-reproducible builds do not inhibit the freedom to study the software! As [XZ Utils](https://en.wikipedia.org/wiki/XZ_Utils_backdoor) taught us, you can always inspect the binaries and check it’s correspondence with the intended code.

Instead, afaik, with some statistical AI systems (LLM, ANN etc…) the only way to ensure that the training data declared by the developers are actually the one used to compute the weights distributed, is to replicate the process.

So the only way to **prevent the Open Source AI definition to becomes a Open Washing AI definition** , used to fool users and get their trust while preventing them to really study the actual system they are using, is to require such reproducibility (that, as said, might amount at worst to recording the whole training process).

I’m more than happy to learn about a different way to obtain the same guarantee about the completeness of the declared training data.

But a OSAID that can be used to open wash black boxes, negating the freedom to study the system and selling the freedom to fine tune as the system to modify, would be [inherently unsecure](https://arxiv.org/abs/2204.06974) and detrimental to users and researchers.

So I hope we can keep brainstorming, looking for a better definition that can really grant the freedoms it aims to grant, including the freedom to study and to modify the system.

---

_[View the full topic](https://discuss.opensource.org/t/draft-v-0-0-9-of-the-open-source-ai-definition-is-available-for-comments/513)._
