Tag: AI Regulation

  • DOJ AI Training Fair Use Brief Hits a 10.8 Million-Work Case

    The DOJ AI training fair use brief, filed September 1, 2026 in the Southern District of New York, tells Judge Sidney H. Stein that training large language models on copyrighted articles is “extraordinarily transformative” and protected. It landed in a consolidated case covering more than 10.8 million asserted works. Three days later, The Seattle Times and Newsday sued OpenAI and Microsoft anyway.

    What does the DOJ AI training fair use brief actually say?

    The Justice Department filed a Statement of Interest on September 1, 2026 in the consolidated OpenAI copyright litigation, MDL 25-md-3143, before Judge Sidney H. Stein in the Southern District of New York. It is not a party to the case. It filed anyway.

    The position is narrow in form and sweeping in effect. Per IPWatchdog’s reading, the DOJ argues that copying works to train a model is fair use because the model converts text into numerical representations rather than using the articles as journalism.

    The brief concentrates on fair use factors one and four: purpose and character of the use, and effect on the market. Associate Attorney General Stanley Woodward Jr. framed it as tying copyright doctrine to US competitiveness in AI.

    The three arguments that matter commercially

    • Transformation. Training copies serve a different function than the original article, so the first factor favors the developer.
    • No market substitution. The DOJ says liability should not turn on anomalous outputs that regurgitate protected text.
    • Oligopoly risk. A licensing mandate would, in the government’s words, “transfer much of the resulting profits to existing publishers with the largest archives” and raise barriers only the biggest technology firms could absorb.

    That third argument is the interesting one. The DOJ is making an antitrust-flavored case for a copyright outcome — arguing that paying for data would concentrate the AI market rather than open it.

    How big is the case the DOJ walked into?

    Very big. Publishers in the consolidated action assert more than 10.8 million works, a figure OpenAI disputes, according to a filing summary published by PPC Land. The New York Times alone asserts 6,030,928 articles published between 1928 and 2024.

    The plaintiff group runs to five news organizations: The New York Times Company, the Daily News plaintiffs covering eight publications, Ziff Davis with six digital properties, the Center for Investigative Reporting, and The Intercept.

    On September 4, 2026 — three days after the DOJ brief — all parties filed summary judgment motions simultaneously. Stein now has to resolve liability and fair use before any trial.

    The damages arithmetic nobody will say out loud

    US statutory damages run from $750 to $30,000 per infringed work, rising to $150,000 for willful infringement. Apply that range to 10.8 million asserted works and the arithmetic is absurd on its face.

    Statutory tierPer work× 10.8M asserted works
    Statutory minimum$750$8.1 billion
    Standard maximum$30,000$324 billion
    Willful maximum$150,000$1.62 trillion
    Anthropic settlement benchmark~$3,000$32.4 billion
    Arithmetic applies the statutory range in 17 U.S.C. §504(c) to the asserted-work count reported in the summary judgment filings. No court has awarded anything close to these sums, and courts routinely group works and reduce awards.

    Treat those numbers as leverage, not forecasts. Even the floor is real money: $8.1 billion is more than most AI companies have ever raised. The Anthropic benchmark is the honest comparison — roughly $3,000 a book across a $1.5 billion settlement approved in 2026, per the Authors Guild. Scale that per-work rate to 10.8 million articles and you get $32.4 billion.

    Why are The Seattle Times and Newsday suing now?

    Because the DOJ brief raised the cost of waiting. The Seattle Times Company and Newsday filed in SDNY on September 4, 2026, alleging OpenAI and Microsoft trained on their journalism without consent and diluted their trademarks by generating fabricated content attributed to them. They seek damages and destruction of the datasets built from their work.

    Seattle Times president Alan Fisco told The Spokesman-Review: “This was not an easy decision… we feel strongly that we must defend our content from being used without our consent or compensation.”

    The awkward detail, noted by TechCrunch: both defendants have previously funded Seattle Times journalism projects and fellowships. Microsoft said it was “surprised by the lawsuit” and “always happy to sit down and explore solutions.” OpenAI said its models are “grounded in fair use.”

    The economics behind the filing

    Search referral traffic to midsize publishers fell 47% year over year in December 2025. When the distribution channel that funded the newsroom collapses, litigation stops being a philosophical position and becomes a revenue line.

    Who profits if the DOJ position wins?

    Model developers, immediately and enormously. If training on copyrighted text is fair use, the largest contingent liability on every frontier lab’s balance sheet evaporates, and the input cost of the next model drops to compute plus salaries.

    OpenAI closed a $122 billion round at an $852 billion valuation in March 2026. A ruling that training data is free removes the one legal event capable of repricing that number downward overnight.

    Publishers lose twice: no damages, and no leverage to price a license. Bloomberg reported Axel Springer’s OpenAI agreement at tens of millions of euros per year — a price negotiated in the shadow of legal risk. Remove the risk and the price goes toward zero.

    The conflict the filing does not mention

    Here is the part that deserves skepticism. In July 2026, OpenAI was reported by the Financial Times and others to be discussing giving the US government a roughly 5% equity stake, a proposal CNBC covered at the time. At an $852 billion valuation, 5% is about $42.6 billion.

    The Justice Department filed a brief advancing OpenAI’s core legal defense while the administration it serves was reportedly negotiating an equity position in OpenAI. The brief does not disclose that. Whatever the merits of the fair use argument — and they are genuinely arguable — the government here is not a disinterested friend of the court.

    Authors Guild CEO Mary Rasenberger called the filing “replete with faulty arguments and a gross misunderstanding of fair use doctrine.” Berkeley copyright scholar Pamela Samuelson called it “a significant development.” Both can be true.

    What does this change for the AI market?

    It shifts the expected value of every US AI copyright case at once. A Statement of Interest does not bind Judge Stein, but it hands every defendant in every pending training-data suit a federal endorsement to cite.

    Timeline of the last two weeks

    DateEvent
    Dec 2023The New York Times sues OpenAI and Microsoft in SDNY
    2026Anthropic’s $1.5 billion author settlement receives final approval
    Jul 2, 2026OpenAI reported to be discussing a ~5% US government equity stake
    Sep 1, 2026DOJ files Statement of Interest backing fair use in MDL 25-md-3143
    Sep 4, 2026All parties file summary judgment motions; Seattle Times and Newsday sue

    Why this matters

    Training data is the last unpriced input in the AI cost stack. Compute is priced — brutally so, as the industry’s multibillion-dollar compute contracts show. Talent is priced. Distribution is priced. Content never has been, and this case decides whether it ever will be.

    For investors, the read is about variance, not direction. Frontier lab valuations embed an assumption that training data stays free — now endorsed by the executive branch and still unresolved by the judiciary.

    Three scenarios are worth holding in mind:

    1. Fair use wins outright. Content costs stay near zero. Publisher licensing revenue collapses. Model economics improve permanently.
    2. Split ruling. Training is fair use, but outputs that reproduce protected text are not. Labs pay for filtering and indemnities, not for data. This is the most likely outcome.
    3. Publishers win on the merits. A per-work price gets set, retroactively, across an industry that has already trained on the corpus. Expect settlements measured in tens of billions.

    Note the second-order effect the DOJ itself flagged. If licensing becomes mandatory, only firms with the balance sheets to pay can build frontier models — the same dynamic visible in consolidation across the AI tooling layer. A publisher win could entrench incumbents rather than dislodge them. The government’s argument is self-serving here. It is not wrong.

    This post is reporting and analysis, not financial advice.

    Frequently asked questions

    Does the DOJ brief decide the case?

    No. A Statement of Interest is advisory. Judge Sidney H. Stein decides the pending summary judgment motions, and any ruling is appealable to the Second Circuit.

    How many works are at issue?

    Publishers assert more than 10.8 million works in the consolidated case. OpenAI disputes the count. The New York Times alone asserts 6,030,928 articles from 1928 to 2024.

    What is OpenAI’s factual defense?

    OpenAI and Microsoft told the court that verbatim reproduction appeared in 24 instances across 20 million chat logs — a rate of 0.00012% — and argued that anomalous outputs should not establish liability.

    How does this compare to the Anthropic settlement?

    Anthropic settled with authors and publishers for $1.5 billion, roughly $3,000 per book. That case involved pirated books rather than licensed news archives, so it sets a settlement benchmark, not a legal precedent. We covered the follow-on exposure in our piece on the music industry’s copyright claims against Anthropic.

    Why does the government’s OpenAI stake matter?

    Because it creates an apparent conflict. OpenAI was reported in July 2026 to be discussing a roughly 5% US government equity stake — about $42.6 billion at its $852 billion valuation — while the DOJ was preparing a brief supporting OpenAI’s central legal argument.

    Will more publishers sue?

    Likely. The Seattle Times and Newsday filings show that the DOJ brief has not deterred plaintiffs. Falling search referral traffic — down 47% year over year for midsize publishers in December 2025 — makes litigation one of the few remaining revenue options.

    What should investors watch next?

    The summary judgment ruling from Judge Stein. It is the first time a US court will resolve fair use for LLM training on a full evidentiary record, and it prices the liability that sits under every AI valuation, including the advertising revenue OpenAI is now building on top of ChatGPT.

    The bottom line

    The DOJ AI training fair use brief is the strongest signal yet that the US executive branch intends to treat training data as free. It does not bind the court, and the same week it landed, two more newspapers sued anyway.

    What happens next is narrow and datable: Judge Stein rules on cross-motions for summary judgment in MDL 25-md-3143. That ruling, not the government’s brief, sets the price of the AI industry’s most valuable unpaid input.

    Expect a split. Expect an appeal. And expect the settlement math — $3,000 a work against 10.8 million works — to shape every licensing negotiation between now and the Second Circuit.

    Sources