Logo
FrontierNews.ai

The Bigger AI Gets, the Harder It Is to Prove What It Copied: Why That's a Legal Nightmare

As artificial intelligence image generators grow larger, they develop what researchers call "attribution decay," making it nearly impossible to prove whether a model copied specific artwork or independently created something similar. MIT computer scientists discovered that the more training data fed into diffusion models like Stable Diffusion and Midjourney, the less traceable their outputs become to any single source image. This finding, published in Nature Communications, upends assumptions about how to regulate and hold AI companies accountable for potential copyright infringement.

What Did MIT Researchers Actually Find?

Zheng Dai and David K. Gifford from MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) set out to solve a practical problem: how to trace an AI-generated image back to the training data that produced it. Their approach involved removing specific pieces of training data from models and testing whether the models could still generate similar outputs. The results were striking.

When researchers removed famous artworks, like Leonardo da Vinci's paintings or the Mona Lisa, from a large model's training data, the model could still reproduce those images and styles. This phenomenon, which they termed "attribution decay," suggests that as models grow larger and ingest more diverse training data, individual pieces of source material become diluted and impossible to isolate.

"Here we show that attribution, characterized as the task of locating a part of the training data that can be held responsible for a generated sample, can become impossible if a model is trained on a sufficiently large corpus of data," the authors stated in their research.

Zheng Dai and David K. Gifford, MIT Computer Science and Artificial Intelligence Laboratory

Why Does This Matter for Copyright and AI Regulation?

The implications are profound for ongoing legal battles. Artists have filed multiple lawsuits against AI companies, including the case Andersen et al. v. Stability AI Ltd, arguing that models were trained on their work without permission and can now reproduce their artistic styles. Proving this requires showing a direct link between the training data and the model's output.

If attribution becomes impossible at scale, the legal foundation for these cases crumbles. Companies could theoretically build models so large that no single output can be definitively traced to any specific training image, creating what amounts to a liability shield through sheer computational size.

David K. Gifford explained the broader implications of the research, noting that the findings raise fundamental questions about intellectual property and compensation in the AI era.

"If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated when what comes out of a model isn't attributable to anything on the internet," Gifford said.

David K. Gifford, MIT Computer Science and Artificial Intelligence Laboratory

How Are Legal Experts Responding to These Findings?

James Grimmelmann, a law professor at Cornell Law School and Cornell Tech, acknowledged that the research fundamentally changes how courts will need to approach AI copyright cases. Current legal strategies rely on proving that a model copied specific works, but if that proof becomes technically impossible, the entire framework shifts.

"If attribution worked, it would reliably tell us whether similarities between a model's output and a copyright-protected work are due to copying or coincidence. But this paper provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying," Grimmelmann stated.

James Grimmelmann, Law Professor at Cornell Law School and Cornell Tech

Grimmelmann noted that existing copyright cases against AI companies have not yet focused specifically on output similarity for image models. German courts have dealt with apparently memorized outputs from music models, and US courts have examined whether training itself constitutes fair use, but the question of whether generated images are derivative works remains largely untested in court.

What Are the Key Implications for AI Companies and Regulators?

The MIT findings create a paradoxical situation for AI regulation and corporate accountability. On one hand, the research suggests that larger models may be more "creative" in the sense that they synthesize information rather than copy it directly. On the other hand, this same characteristic makes it impossible to verify whether companies have actually respected artists' intellectual property rights.

Several critical implications emerge from this research:

  • Attribution Becomes Impossible: As models scale up with more training data, the ability to trace outputs to specific sources disappears entirely, making copyright enforcement nearly impossible.
  • Fair Use Questions Multiply: If outputs cannot be attributed to training data, courts must reconsider whether training on copyrighted material constitutes fair use, potentially reshaping AI development practices.
  • Liability Avoidance Strategy: Companies could theoretically build models large enough that no output can be attributed to any single input, creating a legal gray zone that regulators have not yet addressed.
  • Compensation Models Break Down: If artists cannot prove their work was used, they have no basis for demanding compensation or licensing fees from AI companies.

What Should Stakeholders Know About Attribution Decay?

The concept of attribution decay challenges a fundamental assumption in AI regulation: that larger, more capable models are easier to understand and control. In reality, the opposite may be true. As models absorb more training data, they become black boxes where individual influences cannot be isolated or measured.

Gifford argued that companies should bear the burden of proving their models cannot be attributed to specific sources, rather than placing the burden on artists to prove copying occurred. This would require a shift in how AI companies approach transparency and accountability.

The research suggests that future AI regulation may need to focus less on tracing outputs and more on controlling training data sources, implementing licensing agreements upfront, or requiring companies to demonstrate that their models do not reproduce copyrighted material in recognizable ways. Without such proactive measures, the technical reality of attribution decay means that reactive legal approaches may become obsolete.