Why AI Art Lawsuits May Be Impossible to Win: What a New Study Reveals About Midjourney and Stable Diffusion
Artists suing AI companies for copyright infringement face a fundamental problem: researchers now say it's mathematically impossible to prove which specific artworks an AI model actually used to generate images. A groundbreaking study published in Nature Communications by MIT researchers suggests that even if an AI-generated image looks identical to a human artist's work, there's no way to definitively trace that similarity back to the original artwork in the training data. This finding could reshape how courts handle intellectual property disputes involving generative AI tools like Midjourney and Stable Diffusion.
How Do Diffusion Models Actually Create Images?
To understand why attribution is so difficult, it helps to know how image-generating AI systems work. Diffusion models, the technology powering tools like Midjourney and Stable Diffusion, don't operate like a human artist studying a reference photo. Instead, they learn patterns from millions or billions of training images and use those patterns to generate new images from scratch. The process is so complex that even AI researchers struggle to trace which parts of the training data influenced any given output.
MIT researchers Zheng Dai and David Gifford decided to investigate this mystery by building 24 custom AI models trained on varying amounts of data, from a few hundred images to hundreds of thousands. They then systematically removed individual images from the training data and measured how much the AI's outputs changed. The idea was simple: if removing a specific artwork caused the model to generate noticeably different images, that artwork must have influenced the model's behavior.
What they discovered was surprising and troubling for artists pursuing legal action. The researchers found that as training datasets grew larger, it became increasingly impossible to link any specific output to any specific training image. They call this phenomenon "attribution decay." In other words, the bigger the dataset, the more invisible individual artworks become.
What Does "Attribution Decay" Mean for Artists?
The implications are stark. Commercial diffusion models used by companies like Midjourney are orders of magnitude larger than the test models the MIT researchers used. This means that in real-world AI systems, there is essentially no single human-generated inspiration behind any given image. Even if an AI-generated portrait looks exactly like a famous artist's style, you cannot scientifically prove that the model referenced that artist's work.
The study tested this principle across three different scenarios: general image generation, face generation, and artwork generation. In every case, the result was identical. The researchers noted that "at large training set sizes it not only becomes infeasible to attribute generated images to training images, it also becomes infeasible to attribute generated people to the real people the model was trained on, and to attribute generated artwork to the artists the model was trained on".
"If a diffusion model generates something, you want to be able to say, 'Oh, this part of the training data was responsible,' but our research shows that's not possible at scale," explained Zheng Dai, a fourth-year PhD student at MIT's Computer Science and Artificial Intelligence Laboratory and the study's lead researcher.
Zheng Dai, PhD Student, MIT Computer Science and Artificial Intelligence Laboratory
Why This Matters for Copyright Lawsuits?
Several high-profile lawsuits have been filed against AI companies by artists who claim their work was used without permission to train image generators. These legal cases hinge on proving that specific artworks were part of the training data and directly influenced the AI's outputs. The MIT study suggests this proof may be legally and scientifically impossible to establish.
Attributability is the cornerstone of intellectual property law. To win a copyright case, you typically need to show a causal link between the original work and the infringing copy. But diffusion models don't create that kind of traceable link. The sheer scale of the training dataset obscures the creative process so thoroughly that it's like trying to spot a single tree in the Amazon Rainforest from the International Space Station.
How to Understand the Black Box Problem in AI?
- Scale Obscures Attribution: Larger training datasets make it mathematically impossible to isolate which images influenced specific outputs, even if you remove individual artworks and test the model again.
- Human Creativity vs. Machine Process: When humans create art, they're often consciously aware of their influences and may even reference specific works directly. Diffusion models use the totality of their training data in ways that remain deeply mysterious, even to researchers.
- The Idea-Expression Problem: Copyright law traditionally protects the specific expression of an idea, not the idea itself. But when a machine generates the expression, it's unclear who owns the copyright or whether copyright applies at all.
This research underscores just how alien these AI systems are compared to human creativity. When a painter creates a portrait, they may consciously study their subject or reference other artists' techniques. They can articulate their influences. Diffusion models, by contrast, absorb patterns from millions of images and synthesize them in ways that don't map neatly onto traditional copyright frameworks.
The legal implications extend beyond just image generation. Similar questions about attribution and ownership are emerging across the AI industry, from large language models (LLMs) used for text generation to video synthesis tools. The fundamental challenge remains the same: how do you prove that a machine learning system used a specific piece of training data when the system's internal processes are essentially a black box ?
Despite these challenges, researchers emphasize that understanding how these models work remains crucial for regulation and accountability. "It's important for us to understand how these models work to properly study them or regulate them," Dai noted. The MIT study is a step toward that understanding, even if it raises more questions than it answers about the future of AI copyright law.