California's A.B. 412: Why Mandatory AI Training Disclosures Remain Unworkable
EFF Deeplinks
- California's A.B. 412 is under consideration again, proposing that generative AI developers disclose all copyrighted works used for training.
- EFF opposes the bill, arguing it is practically impossible to implement because no machine-readable database of copyright exists and much online data lacks verifiable ownership.
- The bill risks centralizing power in Big Tech by imposing compliance burdens that only well-funded companies can afford, while potentially stifling startups, non-profits, and independent developers.
- Existing federal law and ongoing court cases regarding fair use already provide mechanisms for copyright holders to address grievances, making state-level regulation redundant and preemptive.
Technical and Practical Hurdles
- There is no unified, machine-readable registry at the U.S. Copyright Office for developers to cross-reference.
- Copyrighted works often lack public samples or clear metadata, particularly with proprietary software or informal online content.
- The requirement to continuously verify massive, unstructured datasets against an incomplete copyright system is considered unworkable.
Economic Impact
- The bill's broad definition of "developer" includes individuals, small organizations, and open-source initiatives.
- Compliance costs will naturally favor large corporations that can afford dedicated legal and compliance teams.
- Stricter regulations may force smaller innovators to exit the market, reducing competition and diversity in AI development.
Legal Context
- Federal courts are currently adjudicating whether AI training constitutes fair use, making it premature for states to impose conflicting or additional regulations.
- The EFF advocates for maintaining copyright governance at the federal level to ensure consistent nationwide standards for both creators and technology developers.