In the latest example of a troubling pattern in the tech industry, Nvidia appears to have scraped troves of copyrighted content for AI training. The $2.4 trillion company reportedly asked workers to download videos from YouTube, Netflix and other datasets to develop commercial AI projects. The training was reportedly to develop models for products like its Omniverse 3D world generator, self-driving car systems and “digital human” efforts.
Source: NVIDIA’s AI team reportedly scraped YouTube, Netflix videos without permission

The new model does distribute royalties according to titles or hours listened, dividing the value of a member’s plan and any additional audiobook credits used “among the titles the member listened to over the course of the month.” Audible’s new royalty model is a response to Spotify’s encroachment into the audiobook space but it is also an early statement to the publishing world about how the company plans to account to publishers and authors.


Given the gaps in existing legal protections, the Office recommends that Congress enact a new federal law that protects all individuals from the knowing distribution of unauthorized digital replicas. The Office also offers recommendations on the elements to be included in crafting such a law.
Negotiations between the tech and news industries over AI have mostly focused on providing data for the broad training of large language models (LLMs) — but now, deal talks are shifting to address narrower use cases, where news publishers may have more leverage.