Why African Creators Are Fighting Stability AI in Court Over Stolen Training Data
African creators are discovering that their work has been scraped into AI training datasets without consent, sparking a wave of copyright lawsuits against companies like Stability AI that could reshape how generative AI is built. The issue goes beyond simple piracy; it represents a systemic appropriation of creative work at a scale never before possible, with profound implications for artists, musicians, and writers across the continent.
How Did AI Companies Build Their Models on Copyrighted Work?
Generative AI systems like Stable Diffusion learn by consuming vast amounts of existing content. Companies including Stability AI and OpenAI, the maker of ChatGPT, built their systems by scraping billions of images, books, articles, and songs from the internet. The process was efficient and cost-effective, but it came at a significant cost to creators whose work was used without permission or compensation.
The LAION-Aesthetics dataset, which was commissioned by Stability AI to train Stable Diffusion, reveals the scale of this appropriation. Studies found that 47 percent of the dataset consists of images from stock photo sites like Shutterstock and Getty, shopping sites including Pinterest, and user-generated content platforms like Flickr. In other words, nearly half of the images used to train Stable Diffusion came from sources that are supposed to require licensing and payment.
Why Does This Matter Specifically for African Creators?
For African creators, this issue strikes at the heart of cultural and economic survival. A photographer in Ghana who uploads images to a stock photo site may find those images have been scraped into an AI training dataset without her knowledge. A musician in Mali whose traditional rhythms are sampled into an AI system may never receive compensation for the use of his cultural heritage. A writer in South Africa whose novel is available online may see an AI system generate derivative works that compete directly with her own book.
The broader context makes this even more urgent. African creators have historically struggled with their work being consumed and enjoyed worldwide while money and recognition flow elsewhere due to piracy, weak copyright enforcement, and lack of infrastructure. The digital age promised to change this through online platforms and streaming services. But just as that promise was materializing, AI companies arrived with a far more sophisticated threat.
What Legal Battles Are Underway?
The resulting lawsuits represent some of the most important legal cases of our time. One of the most significant began in January 2023, when visual artists filed a class action lawsuit in the United States against Stability AI, Midjourney, and DeviantArt. The lead plaintiffs, Sarah Andersen, Kelly McKernan, and Karla Ortiz, were joined by other artists including Jingna Zhang, Gerald Brom, and Greg Rutkowski, alleging that these companies had used their copyrighted works to train AI image-generation models without permission.
The case, known as Andersen v. Stability AI, has progressed through the American legal system with significant developments. In August 2024, U.S. District Judge William Orrick issued a ruling allowing the artists to pursue claims that the AI image generators infringe upon their copyrights. The judge found that the artists had reasonably argued that the companies violate their rights by illegally storing their work and that Stable Diffusion "may have been built 'to a significant extent on copyrighted works' and was 'created to facilitate that infringement by design'".
Meanwhile, Getty Images, one of the world's largest stock photography companies, launched separate proceedings in the High Court in London against Stability AI in 2023, originally alleging infringement of copyright in millions of images used to train Stable Diffusion, as well as infringement of database rights and trademarks. The case pushed into uncharted legal territory on the use of copyright materials in AI model training.
Steps Creators Can Take to Protect Their Work
- Monitor AI Training Datasets: Creators should research whether their work appears in publicly documented AI training datasets like LAION-Aesthetics and document any unauthorized use for potential legal action.
- Join Collective Legal Action: Artists can explore joining class action lawsuits or collective licensing agreements that pool resources and increase leverage against large AI companies.
- Assert Data Sovereignty Rights: Advocate for policies that establish African creative data sovereignty, ensuring that African people retain control over how their cultural expressions are used in AI systems.
- Use Licensing and Attribution Tools: Implement clear licensing terms on published work and use metadata to track how images and content are distributed across the internet.
- Support Regional Legal Frameworks: Engage with policymakers to develop stronger copyright enforcement mechanisms and AI regulation specific to African creative industries.
What Happened in the Getty Images Case?
The Getty Images case demonstrates the complexity of applying existing copyright law to AI. As proceedings progressed, Getty narrowed the scope of its claims, ultimately deciding not to pursue its claim of primary infringement. The central challenge was that training had taken place entirely outside the UK and thus principally outside the reach of UK copyright law.
Instead, Getty rested its copyright claim on whether offering access to Stable Diffusion in the UK constituted the importation of an infringing article. The High Court judgment held that while an "infringing article" may consist of intangible property, the model weights underpinning Stable Diffusion did not reproduce Getty's images and so did not constitute an infringing article.
However, Getty Images was granted permission to appeal. In granting permission, Mrs Justice Joanna Smith acknowledged that the Court of Appeal may arrive at a different conclusion and that Getty's proposed appeal "does have a real prospect of success. It concerns a pure question of law, namely a matter of statutory construction on which the minds of reasonable lawyers may differ". The judge also recognized the broader importance of the case, noting that "the point of law is both novel and important because it concerns how the provisions of the CDPA should be construed (and specifically the phrase 'infringing copy') in the context of an AI model".
In the United States, Getty Images achieved a partial victory in its separate lawsuit against Stability AI when a U.S. District Court in California found that Getty Images sufficiently alleged claims for trademark infringement and unfair competition against the company.
What's at Stake for the Creative Economy?
These legal battles will likely reshape how AI companies build their models and how creators are compensated for their work. The outcomes could establish precedents for whether AI training on copyrighted material constitutes infringement, how "infringing copies" are defined in the context of AI models, and what obligations companies have to compensate creators whose work was used without permission.
For African creators specifically, the stakes involve not just individual compensation but the preservation of cultural sovereignty. As the continent's creative industries continue to mature, disputes surrounding originality, attribution, ownership, and control of creative works are becoming more frequent and legally sophisticated, particularly in an era driven by digital platforms, global distribution, AI-assisted creation, and viral social media dissemination. The creator economy in Africa is not just an economic opportunity; it is a matter of cultural survival.