MIT Researchers Find AI Models Lose Source Attribution with Scale, Complicating Regulation
Researchers at MIT's Computer Science & Artificial Intelligence Laboratory (CSAIL), Zheng Dai and David K Gifford, have published a paper titled "Outputs of Generative Diffusion Models are Often Unattributable," which will appear in Nature Communications. Their findings indicate that as generative diffusion models, such as Midjourney and Stable Diffusion, grow larger and are trained on more data, their outputs become increasingly difficult to attribute to specific source material. This phenomenon, termed 'attribution decay,' suggests that the more data a model processes, the less it 'remembers' the origin of its generated content. The researchers tested this by removing specific training data, like images of the Mona Lisa or all of Leonardo Da Vinci's work, and found that very large models could still reproduce similar images or styles. This challenges the notion that AI models merely copy their training data, suggesting a form of creativity.