Eigen RadarAI
Analysis

Consent and openness show two paths for AI data

Twitch is offering an opt-out for stream recordings while OpenWALDO builds an open, contributed corpus; both developments highlight consent, origin and usage information in access to training data.

Artificial Intelligence··Morning
In a dark broadcast studio, an anonymous creator closes an amber physical lever that stops two blue carriers in front of a camera and microphone.

Default use on Twitch, followed by an opt-out

Twitch has added a setting that allows streamers to stop their recordings from being used to train Amazon's generative AI models. The toggle appears in the generative AI training section under security and privacy, and it is enabled by default, so a streamer must change the setting to halt use. Chief product officer Mike Minton told TechCrunch that the company chose this arrangement because it believed nobody would opt in if participation were voluntary. Twitch presented the change as the addition of an opt-out control rather than as the announcement of a new training programme. Minton also said he did not know whether creator content had already been used for training. The report notes that the platform holds thousands of hours of streamer audio and video and that much of its community opposes generative AI because of concerns about unauthorised scraping. The arrangement leaves an access decision to individual streamers in their channel settings, but the default position makes a clear distinction between silence and affirmative consent.[1]

OpenWALDO seeks to grow data through contribution

Gregory Kurtzer, the founder of CentOS and Rocky Linux, has launched OpenWALDO, an AI training dataset funded by his company CIQ. The openly contributed collection contains 167.3 billion reference tokens drawn from government records, open academic papers, mailing lists and public-domain literature. Its name expands to Open Weights, Artifacts, Licenses, Data, Origins. The project asks contributors to grow the corpus in the way open-source software is built and aims to establish its foundations in the open; it is published on its own website and on GitHub. The Register notes that the volume remains small beside the trillions of tokens used by frontier developers for pretraining. CIQ did not answer whether any model has yet been trained on the dataset. The launch therefore supplies information about the origin and contribution model of training material rather than performance results for a finished model. Naming the source categories, placing licences and origins in the project's name, and inviting contributions through GitHub make parts of the collection process visible. The report does not yet contain measured results showing what model behaviour the corpus produces. OpenWALDO's concrete scale today is 167.3 billion reference tokens; its role in larger training programmes would require further records of actual use.[2]

Access rules and origin records fill different gaps

Twitch and OpenWALDO occupy different points in the training-data lifecycle. Twitch holds audio and video recordings produced by streamers on its platform, and the new control lets each streamer stop those recordings from being used for Amazon models. OpenWALDO gathers material from four source categories in a contributed dataset and makes origins, licences and artifacts part of the project's identity. In the first case, the central access question is whether a creator must act to halt use. In the second, it is how a collection with named origins is assembled and expanded. The sources do not say that the two initiatives follow the same consent system or operate at the same scale. The Twitch report leaves earlier training use unknown, while the OpenWALDO report leaves unanswered whether a model has yet been trained on the collection. Each development therefore supplies a visible control or data description without providing the complete history of actual use. Their shared context is that access to training data cannot be described by a count of tokens or hours alone. Who supplied the material, the rule under which it may be used and where it has actually been used are separate pieces of information. Twitch foregrounds the consent side of the first two questions. OpenWALDO foregrounds source origins and contribution. Actual model use remains open in both reports.[1], [2]

References

  1. News sourceTechCrunchTwitch will train Amazon's generative AI models on stream recordings unless creators opt out↩1↩2
  2. News sourceThe RegisterRocky Linux's founder launches OpenWALDO, an open training dataset of 167.3 billion tokens↩1↩2