AI copyright regulation is not a side quest. For many model providers, it is the unresolved line item that sits under every product launch and every partnership. U.S. AI policy news has already produced lawsuits, Copyright Office reports, licensing experiments, and a lot of speeches about “balance.” What it has not produced, yet, is a stable public answer to a few questions companies would prefer to keep in the appendix.
I am not here to cheer for either maximal permission or maximal scraping. I am here to name the questions that keep being deferred because answering them would force a business-model conversation.
The Questions That Will Not Stay in the Footnotes
Did the training corpus include copyrighted works at scale? Under what legal theory—fair use, license, implied permission, or a bet that plaintiffs cannot prove the copies? If outputs sometimes reproduce protected expression, is that a bug, a statistical inevitability, or a product feature for some users? If a license is required, who pays, at what rate, and for which vintages of data?
Those are not philosophical puzzles. They are operational. They affect what can be trained, what can be sold to enterprises that fear indemnification gaps, and what publishers will block.
The deferral pattern
Speak about “publicly available” data as if availability were the same as a license.
Point to a transformation argument without specifying how much transformation is enough.
Offer opt-outs after the fact, which do not unwind a completed training run.
Market enterprise indemnification that is capped, conditioned, or silent on training.
Wait for a court in a specific circuit rather than setting a public standard.
Read the announcement. Then read the incentives. The incentive is to ship the model now and let the legal theory catch up.

What Enterprises Are Actually Buying
When a company sells an API or a copilot, the customer is buying a capability and a risk allocation. If the vendor’s training practices are opaque, the customer is taking a risk they cannot underwrite. Some large vendors now offer indemnity for certain infringement claims arising from outputs. The details matter: whose prompts, which products, which jurisdictions, which excluded uses.
Publishers, newsrooms, image libraries, and software-rightsholders have taken different paths: sue, license, block crawlers, or do all three. None of those paths is a full market yet. A handful of high-profile licenses do not tell you the clearing price for the long tail of works.
A map of unresolved legal objects
Object | Why it matters | What is still fuzzy |
|---|---|---|
Training copies | Were they made, stored, and used? | Status under fair use in this context |
Output similarity | When is an output a copy, a style, or a coincidence? | Tests courts will actually apply |
Robots.txt and terms | Do site rules bind model trainers? | Technical and legal enforceability |
Opt-out registries | Can a rightsholder exit future runs? | Retroactivity and completeness |
Indemnity | Who pays if a customer is sued? | Caps, carve-outs, and proof burdens |
AI copyright regulation debates that stay at the level of “artists versus innovation” skip this table. The table is where counsel lives.

Why Model Size Became a Distraction
Labs like to talk about scale because scale is a capability story. Rightsholders like to talk about scale because it sounds like industrial copying. Both can be true, and still miss the point that a smaller, domain-specific model trained on a dirty corpus can create as much legal exposure as a famous large one. The relevant facts are provenance, memorization, and use—not parameter counts as a morality score.
I have watched product people try to “solve copyright” with a filter that reduces verbatim regurgitation. Filters help. They do not answer whether the training copy was lawful. They also do not help a newsroom that does not want its archives used as unpaid raw material even if the outputs never quote a paragraph.
Questions I now ask in every briefing that touches data
What is the provenance story you are willing to put in a contract?
Which categories of works are licensed, excluded, or simply unmentioned?
How do you handle a documented regurgitation incident in production?
What does your indemnity exclude?
If a court disagrees with your fair-use theory, what is the product plan?
If the answer is that the issue is “being monitored,” you have been told that the business still depends on a legal outcome someone else will write.
What a Serious Public Settlement Would Have to Include
A durable settlement—whether through Congress, courts, or markets—would need a way to license at scale, a way to exclude, a way to audit claims of exclusion, and a way to handle past training runs that cannot be magically untrained. That is a lot of plumbing. Plumbing is not a keynote topic. It is why the questions keep being avoided.
Here is what changed, and what did not. Generative products became useful enough that rightsholders could measure harm and users could measure value. The underlying copies-and-outputs problem did not dissolve. Who really benefits, and who really pays? Firms that trained early on broad crawls benefit from delay. Writers, photographers, and publishers pay in unconsented use unless a license market or a judgment arrives. Users pay if products shrink or prices jump when the bill comes due.
Policy Desk will keep putting the legal objects on the table. The speeches can wait.
No notes on this sheet yet.