The Whitfield Brief
Policy Desk

Copyright, Training Data, and the Legal Questions Companies Keep Avoiding

Copyright, Training Data, and the Legal Questions Companies Keep Avoiding
This Policy Desk essay names the copyright and training-data questions AI companies often defer: provenance, fair use, outputs, opt-outs, and indemnity. It maps unresolved legal objects and lists briefing questions for enterprises.

AI copyright regulation is not a side quest. For many model providers, it is the unresolved line item that sits under every product launch and every partnership. U.S. AI policy news has already produced lawsuits, Copyright Office reports, licensing experiments, and a lot of speeches about “balance.” What it has not produced, yet, is a stable public answer to a few questions companies would prefer to keep in the appendix.

I am not here to cheer for either maximal permission or maximal scraping. I am here to name the questions that keep being deferred because answering them would force a business-model conversation.

The Questions That Will Not Stay in the Footnotes

Did the training corpus include copyrighted works at scale? Under what legal theory—fair use, license, implied permission, or a bet that plaintiffs cannot prove the copies? If outputs sometimes reproduce protected expression, is that a bug, a statistical inevitability, or a product feature for some users? If a license is required, who pays, at what rate, and for which vintages of data?

Those are not philosophical puzzles. They are operational. They affect what can be trained, what can be sold to enterprises that fear indemnification gaps, and what publishers will block.

The deferral pattern

  • Speak about “publicly available” data as if availability were the same as a license.

  • Point to a transformation argument without specifying how much transformation is enough.

  • Offer opt-outs after the fact, which do not unwind a completed training run.

  • Market enterprise indemnification that is capped, conditioned, or silent on training.

  • Wait for a court in a specific circuit rather than setting a public standard.

Read the announcement. Then read the incentives. The incentive is to ship the model now and let the legal theory catch up.

Contracts at the center of U.S. AI policy news on training data

What Enterprises Are Actually Buying

When a company sells an API or a copilot, the customer is buying a capability and a risk allocation. If the vendor’s training practices are opaque, the customer is taking a risk they cannot underwrite. Some large vendors now offer indemnity for certain infringement claims arising from outputs. The details matter: whose prompts, which products, which jurisdictions, which excluded uses.

Publishers, newsrooms, image libraries, and software-rightsholders have taken different paths: sue, license, block crawlers, or do all three. None of those paths is a full market yet. A handful of high-profile licenses do not tell you the clearing price for the long tail of works.

A map of unresolved legal objects

Object

Why it matters

What is still fuzzy

Training copies

Were they made, stored, and used?

Status under fair use in this context

Output similarity

When is an output a copy, a style, or a coincidence?

Tests courts will actually apply

Robots.txt and terms

Do site rules bind model trainers?

Technical and legal enforceability

Opt-out registries

Can a rightsholder exit future runs?

Retroactivity and completeness

Indemnity

Who pays if a customer is sued?

Caps, carve-outs, and proof burdens

AI copyright regulation debates that stay at the level of “artists versus innovation” skip this table. The table is where counsel lives.

Why delayed opt-outs do not settle AI copyright regulation

Why Model Size Became a Distraction

Labs like to talk about scale because scale is a capability story. Rightsholders like to talk about scale because it sounds like industrial copying. Both can be true, and still miss the point that a smaller, domain-specific model trained on a dirty corpus can create as much legal exposure as a famous large one. The relevant facts are provenance, memorization, and use—not parameter counts as a morality score.

I have watched product people try to “solve copyright” with a filter that reduces verbatim regurgitation. Filters help. They do not answer whether the training copy was lawful. They also do not help a newsroom that does not want its archives used as unpaid raw material even if the outputs never quote a paragraph.

Questions I now ask in every briefing that touches data

  1. What is the provenance story you are willing to put in a contract?

  2. Which categories of works are licensed, excluded, or simply unmentioned?

  3. How do you handle a documented regurgitation incident in production?

  4. What does your indemnity exclude?

  5. If a court disagrees with your fair-use theory, what is the product plan?

If the answer is that the issue is “being monitored,” you have been told that the business still depends on a legal outcome someone else will write.

What a Serious Public Settlement Would Have to Include

A durable settlement—whether through Congress, courts, or markets—would need a way to license at scale, a way to exclude, a way to audit claims of exclusion, and a way to handle past training runs that cannot be magically untrained. That is a lot of plumbing. Plumbing is not a keynote topic. It is why the questions keep being avoided.

Here is what changed, and what did not. Generative products became useful enough that rightsholders could measure harm and users could measure value. The underlying copies-and-outputs problem did not dissolve. Who really benefits, and who really pays? Firms that trained early on broad crawls benefit from delay. Writers, photographers, and publishers pay in unconsented use unless a license market or a judgment arrives. Users pay if products shrink or prices jump when the bill comes due.

Policy Desk will keep putting the legal objects on the table. The speeches can wait.

Revised · 2026-09-19 10:50
Margin Notes

No notes on this sheet yet.

Add a Note
© 2026 The Whitfield Brief. All rights reserved. drawn by hand