Elizabeth Lyon, an Oregon-based author, recently filed a class-action lawsuit against Adobe, accusing the company of using an illegal dataset containing a large number of pirated books to pre-train its small language model, SlimLM. The accusation strikes at the heart of the AI industry’s core pain point: the legality of training data. As a key component of Adobe’s AI strategy, SlimLM is optimized for document-assistance tasks on mobile devices. However, the plaintiff argues that the open-source dataset SlimPajama-627B, on which the model relies, is a derivative of the RedPajama project. RedPajama has been controversial for including the "Books3" database, which contains 191,000 copyrighted books.\n\nAt the center of the lawsuit is Lyon’s claim that several of her nonfiction writing guides were included in the training data without authorization, attribution, or compensation, directly violating fundamental copyright principles.
文章图片 2
Notably, Adobe has previously promoted its AI products, such as Firefly, as being built on legally sourced and protected content. The allegations against SlimLM now expose potential compliance risks in the underlying technology, posing a significant threat to Adobe’s brand image and market trust.\n\nThe AI industry is experiencing a legal storm driven by data compliance issues. Anthropic, for example, previously paid a staggering $1.5 billion in damages, serving as a stark warning to the entire sector. As global regulators tighten oversight of AI technology, companies must rethink how they acquire and use training data. Against a backdrop of intensifying competition in computing power, balancing technological innovation with copyright protection has become a strategic challenge for AI enterprises. Training large language models requires not only powerful chips and infrastructure but also a lawful and compliant data foundation—a factor directly tied to the sustainable development of the AI industry.