HomeLearnCoursesHackathonsAccount
AI Governance & Regulation
Recurring Debate Two: Copyright and Liability · 1/2

Training data and copyright: genuinely unresolved

The third recurring debate, and one of the least settled, is whether training a model on copyrighted material without a license constitutes infringement, and if so, under what circumstances. This is not a question with a confident answer to give right now: it is the subject of active litigation in multiple jurisdictions, and courts have not converged on a single doctrine. In the US, much of the debate centers on whether training qualifies as 'fair use,' a flexible, fact-specific defense to copyright infringement that weighs factors like the purpose of the use and its effect on the market for the original work. Different courts have reached different preliminary conclusions on different facts, and none of this should be read as a settled rule a builder can rely on.

What can be said with more confidence is the shape of the disagreement, not its outcome. Rights holders argue that training on their copyrighted work without permission or compensation is exactly the kind of commercial use copyright law is meant to require licensing for, especially when the resulting model can be used to produce outputs that compete with the original creators' work. Model developers have argued that training is transformative, extracting statistical patterns rather than reproducing protected expression, which they contend places it closer to established fair-use precedents around indexing and analysis. Both positions have genuine legal grounding, and that's precisely why this remains unresolved rather than a case of one side being obviously right. A builder relying on a third-party foundation model should treat this as an active legal risk to monitor, not a settled question to assume away in either direction.