MIT finds diffusion models 'forget' source data as they scale
TL;DR. MIT computer scientists found that diffusion models lose direct attribution to source material with increased training data, complicating regulation. - The research by Zheng Dai and David K Gifford from CSAIL challenges assumptions about model memory. - This 'convenient amnesia' makes tracing output back to specific inputs more difficult. - Findings suggest current methods for intellectual property and regulation may need re-evaluation.
- MIT research shows diffusion models become less attributable to original training data as they grow.
- This 'amnesia' makes it harder to link generated outputs to specific source materials.
- The findings complicate efforts to regulate AI and enforce intellectual property rights.
- Paper titled 'Outputs of Generative Diffusion Models are Often Unattributable' will be published in Nature Communications.