
Small business websites relying on organic search traffic should monitor how AI answer engines impact their visitor acquisition channels.
What Do The Unsealed Microsoft OpenAI Filings Say?
Newly unredacted filings in the 3-year-old New York Times copyright lawsuit against OpenAI and Microsoft, originally filed in December 2023, show internal documents describing AI scraping as theft. A top Microsoft executive described the training practices as theft in internal communications, and OpenAI leadership called its models an existential threat to the publishers whose work trained them.
The material details how the content was obtained: bypassing paywalls, building datasets through mass scraping, and stripping copyright notices from training data before it reached the models. The reporting carries one caveat worth keeping: much of the new material comes from the Times’ own brief, and the underlying exhibits remain sealed.
A January 2023 internal memo by Microsoft’s Brent Hecht called the practice an astonishing theft of unprecedented proportions and the largest theft of labor in human history.
The filings turn a legal argument into a documented account of how answer engines were built and what they cost publishers.
Does The 93% Copilot Click Drop Hold Up?
The figure comes from Microsoft’s own data, cited in the newly unredacted material: Copilot’s answer engine cut click-through rates to the New York Times domain by as much as 93% compared with traditional Bing search. That is the platform measuring its own substitution effect.
A January 2024 presentation by Brent Hecht, Microsoft’s director of Applied Science, called the decline a doom loop that would hurt the performance of their models and the entire web at the same time. The document states it is highly unusual that an end-product threatens the economic foundations of its essential suppliers.
OpenAI’s head of ChatGPT, Nick Turley, wrote that publishers face an existential threat from products that are “largely substitutive” and “will get more and more substitutive as they get better”. Satya Nadella agreed under oath that chatbots substitute giving you the information right there versus visiting the underlying source.
The number is the platform’s own measurement, which makes it harder to dismiss than any publisher estimate.
How Is AI Search Different From The Old Search Bargain?
Traditional search traded a listing for a click: the engine indexed your page, the visitor clicked through, and you monetized the visit on your terms. An answer engine reads your content and keeps the visitor on its own platform.
The scale documented in the filings is industrial. OpenAI’s mid-training datasets contain more than 91,692 copies of works from the NYT, Daily News and Center for Investigative Reporting, and a dataset derived from Common Crawl, the open repository of web crawl data, included more than 2 million documents from nytimes.com alone.
Project Mango, a Microsoft and OpenAI data initiative, produced a training dataset with copies of at least 160,903 unique works from news publishers. OpenAI also delivered the entire GPT-3 training dataset to Microsoft for product evaluation, and the case history is laid out on the public record page for the case.
Every layer of the new stack, from the crawl to the answer, was built to hold the visit.
Which Sites Lose The Most Traffic To AI Search?
Informational content carries the exposure. Any business that publishes answers for free to earn the visit now competes with the platform answering those questions itself, and the filings show the substitution effect is intentional and improving.
Commercial and transactional pages still collect visits because the user needs to act, not read. That split, informational pages feeding answers while commercial pages collect clicks, is where a small business should read its own exposure.
The pages that still earn clicks are the ones worth optimizing hardest, and the discipline behind that work is what our NeuronWriter intelligence report covers in depth.
The visitor center at the interstate exit answers the questions travelers used to drive into town for, with a laminated sheet and a photo of the diner 2 miles in. Travelers read the summary, nod, and merge back onto the highway. The diner feeds the answers, and the visitor center keeps the visit.
Microsoft’s own number for that counter is a 93% cut in click-through on the Times domain, and it comes from the platform’s own measurement. Hecht’s January 2024 slide called it a doom loop that hurts the models and the entire web at the same time, because the product starves the suppliers it needs.
Your informational pages are the diner’s photo now: they feed answers that end the trip before it starts. The commercial pages, the ones where the visitor must act, still collect the visit, and that split is where the budget moves.
Informational content is brand spend from this quarter on, and acquisition budget belongs where the click still lands.
What Should You Do About AI Search Traffic Now?
Pull 90 days of organic sessions and split them by intent. Informational pages feed the answers, commercial and transactional pages still collect the visit, and the split tells you which pages exist for brand rather than acquisition.
Shift acquisition spend toward channels an answer engine cannot intercept: email, direct relationships, communities, and partnerships. Own the relationship, because the engine already owns the query.
Watch the licensing thread in the case. Nadella testified that anything paywalled should be licensed by anyone who wants to use it, which signals where platform economics may move next.
Treat informational content as brand spend this quarter, and put acquisition budget where the click still lands.
Source: TechCrunch AI