Web Analytics Made Easy - Statcounter

Tracking trends in AI policy across open source communities

Policy is one engine through which OSS communities can influence model builders and the AI economy. Collectively stating and enforcing values has huge potential to shift the power balance.

Tracking trends in AI policy across open source communities
Photo by Amélie Mourichon / Unsplash

In August, the CHAOSS AI Alignment Working Group shipped two candidates one of those was a metric model: Governed AI use in Open Source Communities. We will finalize and publish a bit later in September (keep an eye out!).

In the meantime, I wanted to test one of the metrics in Governed AI Use: consent policy specificity, which pulls together a taxonomy of the policy themes we have observed across the policies we have documented over the past year.

My first pass was simply to look for matches, or partial matches (leans), within each policy. What turned up was interesting as much for the expected (code contribution rules) as the unexpected: no rules at all about environment or infrastructure, and not a single mention of notetaker bots, which comes up all the time in my community conversations.

Horizontal bar chart of AI policy coverage across 39 open source projects and nine domains. Code contributions 25 addressed and 13 partial; Content and documentation 6 addressed and 13 partial; Moderation and enforcement 0 addressed and 13 partial; Review of contributions 4 addressed and 10 partial; Autonomous and agentic use 8 addressed and 11 partial; Data use for training 0 addressed and 1 partial. No project addresses Notetaker and meeting bots, Environmental impact or Infrastructure strain.
Analysis of policy documents by metric taxonomy/theme

When I shifted to more broadly analyze policy preamble, blog posts (rationale) it was clearer that considerations like the Environment DID weigh the decisions, but (unless AI was banned) have not specifically made it into policy rules itself.

Horizontal bar chart. Counts of projects that raise a domain in their stated reasons without governing it, out of 39: Environmental impact 7, Infrastructure strain 5, Review of contributions 4, Data use for training 2, Autonomous and agentic use 1. Other projects may still govern these domains; the count is of policies that name the concern and set no rule about it.
Analysis of policy justification/rationale as found in preamble/blog posts etc

I do think our taxonomy will need to expand, based on things like 'openness' of model/data etc used, but this is a start!

The full report, linking to individual project reports can be found in my repo. As you can see I used Claude to help analyze policy, but that each finding/report includes the quote associated with a finding - and I have reviewed and polished each one. There may be some discrepancies in whether something appeared in a policy, or preamble - but the margin of error should be small, and since I remain unpaid for this work - part of the deal for now.

I believe, personally, that policy is one engine through which open source communities can influence model builders and the economy of AI. By stating and enforcing values collectively, we have power. I hope to see projects move from anecdotal statements of frustration and distrust to policy that drives change, especially around infrastructure strain, data use for training and environmental impact, as well as day-to-day encounters (ARGH, note-taker bots).

I know it's a huge challenge, and that these are baby-steps, but capturing as trends like this is important to show that we are not alone, that values of communities creating code and content matter for the future, and not just the past that AI was trained on.

Please consider becoming a paid subscriber, to support my work.

Other Reading

Licensed under CC BY-SA 4.0