AI-Generated Images Prove Impossible to Attribute to Human Creators

·

A new study has shed light on the workings of image-generating AI models, revealing that it’s impossible to determine which human images or artworks were referenced in creating an AI-generated output. The research, published last week in Nature Communications, used a novel approach to understand how diffusion models arrive at their final outputs: by taking them apart and examining each component individually.

The study was led by Zheng Dai, a fourth-year PhD student at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), who wanted to know whether AI-generated images reference specific human artworks or faces. To answer this question, Dai and his coauthor, David Gifford, built custom models trained on online image databases, ranging from hundreds of images to hundreds of thousands.

The researchers then ran a series of experiments to see how the models’ outputs would change if they removed particular images from their training datasets. They found that as the number of images in the dataset increased, it became increasingly difficult to determine which specific images influenced the AI-generated output. This phenomenon is dubbed ‘attribution decay.’

In essence, the study demonstrates that at large training set sizes, it’s not only impossible to attribute generated images to their human counterparts but also to identify the real people or artists behind them. The sheer scale of the training dataset obscures its creative process, making it akin to trying to spot a single tree in the Amazon Rainforest from space.

This has significant implications for artists who try to sue AI companies on grounds of intellectual property theft. No matter how much an AI-generated image may resemble their own work, they cannot be certain which images were referenced by the model to create its output. As Dai notes, commercial diffusion models are many orders of magnitude larger and thus far less attributable than the test models used in the study.

The researchers also highlight that these black box systems operate differently from human creators. When humans produce art or writing, they’re usually aware of the influences shaping their work. In contrast, AI-generated images rely on the totality of their training data to create new outputs, often in ways that remain mysterious even to us.

Dai emphasizes the importance of understanding how these models function for proper study and regulation: ‘It’s essential for us to comprehend how they operate so we can properly regulate them.’ The study underscores just how far legal experts still have to go to untangle the complex intellectual property questions surrounding AI-generated images.