Wednesday, September 16, 2026

Project Syndicate -- AI Is Fracking the Open Web -- Sep 14, 2026 -- Lucky Gunasekara

 Project Syndicate

AI Is Fracking the Open Web

Sep 14, 2026

Lucky Gunasekara

The demand for what digital publishers offer is real and robust, but the old methods of managing it will no longer work when cyberspace is inundated with bots designed to extract value wherever it can be found. Building an alternative will fall to the remaining advocates of the old open web.



Listen to this commentary

1x

BOSTON—At some point in the AI race, the internet came to be thought of as a coal mine. To those building frontier models with piles of GPUs and venture capital to burn, the open web is a vast raw seam of informational material just waiting for someone with enough computing power and cleverness to release its potential. Just send in the crawlers, haul out the data, and feed it all into the furnace of machine intelligence.


What Amodei’s Call for an AI Pause Gets Wrong

Innovation

What Amodei’s Call for an AI Pause Gets Wrong

Gabriela Ramos thinks the Anthropic CEO’s appeal for a slowdown in model development leaves too many unanswered questions.


China’s External Surplus Makes a “Mandarin Accord” Inevitable

Economics

China’s External Surplus Makes a “Mandarin Accord” Inevitable

Jim O'Neill sees a growing need for international coordination to address growing global imbalances.


Why India Wants Peace in Ukraine—and China Doesn’t

Politics

Why India Wants Peace in Ukraine—and China Doesn’t

Nina L. Khrushcheva observes that the longer the war drags on, the weaker and more dependant Russia becomes.




But this mental model glosses over several points. After all, mines have concentrated deposits, clear owners, leases, and downstream refineries, whereas the open web has none of these things. Rather than mining the internet, AI is fracking it, plumbing every nook and cranny to see what previously unavailable resources might be extracted. Valuable information is scattered across pages, PDFs, feeds, archives, application programming interfaces (APIs), paywalls, forums, social-media posts, audio and video libraries, proprietary datasets, code repositories, and a vast field of irregularly maintained HTML pages. And so, the AI industry drills wherever it can.


Bots fetch public pages, scrapers work around brittle defenses, and “headless browsers” emulate the behavior of human visitors. Residential proxy networks make automated scraping look like page visits from your next-door neighbor. Even search results become back doors for data harvesting. And atop this apparatus of digital pumps and hoses sit frontier labs, AI search and RAG (retrieval-augmented generation) engines, retrieval services, data brokers, and AI agents, all trying to extract the same thing: not raw HTML, but the facts, claims, quotes, prices, timelines, source trails, judgments, and expert opinions buried within.


This extractive swarm clearly represents a crisis for the open web. But it also points to an uncomfortable truth: extraction proves demand. The only question is: Who gets to own and control the value at stake? The web isn’t a natural resource; it is a product of human labor. It is a collective of reporters breaking investigations, researchers sharing hard-won discoveries, developers patching open-source software, and artists sharing new works. If the incentives to keep doing this collapse, new contributions to the commons will inevitably run dry.


The Rise of Machine Demand

There is no clearer sign of this industrialized digital fracking than the rapid shift in online traffic from humans to bots and crawlers. Multiple infrastructure providers now report that bots account for roughly half or more of observed internet traffic. Not all of that is necessarily AI-related, but the larger shift is clear. Machines are no longer in the minority online; we are.


Make your inbox smarter.

Select Newsletters

For now, traffic to news sites is hovering around 63% human. Readers have not disappeared, at least not yet. But they are increasingly served by machine interfaces that read, retrieve, summarize, rank, cite, and act on their behalf. The emerging agentic web is therefore WALL-E-like. Why lift a finger to click a source when a machine will consume and pre-digest it for you? Pew Research Center found that when Google displayed an AI summary, users clicked a traditional search result only 8% of the time, compared with 15% when no summary appeared. Source links inside the AI overviews were clicked in only about 1% of visits.



PS Quarterly: Slopocalypse Now is here. As a Premium subscriber, you have exclusive access.


AI-generated content already accounts for half the articles published online, and not even digital forensics experts can tell what’s real. Can those still offering a humanistic vision of online life make themselves heard over the din?


READ NOW


But there is an important distinction to make between crawlers (for AI training) and “fetchers” (for everyday AI queries). Training crawlers want snapshots of the past. They need to gather archives and vast collections of data to build future models. By contrast, user-action fetchers, AI search bots, RAG engines, and agents want the present. They scour the web in real time to answer a question directly, generate a summary, verify a claim, monitor an event, or complete a task.


The distinction matters because each represents a different form of demand. Training bots reflect AI firms’ bottomless appetite for bulk data. Fetchers and agents are more reflective of the demand from consumers, researchers, and professionals seeking timely, authoritative information.


The fingerprints of this demand are becoming widely visible. Cloudflare found that user-action crawlers followed daily usage cycles, with ChatGPT driving roughly three-quarters of requests. HUMAN Security likewise found that inference-time scraping for RAG, live search, news summaries, pricing, and comparison tools grew 597% in 2025, while direct agentic traffic grew 7,851% from a small base.


Fastly likewise found that AI crawlers accounted for roughly 80% of AI-bot traffic. But in a later publisher study, the distribution flipped: fetchers accounted for 82% of observed publisher AI traffic, and crawlers for only 18%. TollBit similarly found that RAG bots requested publisher pages about ten times more than training bots. A training crawler may visit once to take a snapshot, but a RAG bot returns repeatedly because the answer must remain current.


Shuwei Fang of Harvard’s Shorenstein Center describes this demand-side shift as the rise of machine audiences, and with it a movement from attention to intention. Given the nature of this change, we shouldn’t confuse illegitimate extraction with illegitimate demand.


Crumbling Norms and Defenses

Crawler operators (and their lawyers) will argue that their work is transparent and respects robots.txt, the plain-text web standard stating which bots may access a website, and which may not. But while the web has a consent file, it does not have the means to enforce it, and it is precisely here that permitted forms of extraction break down and give way to a black (or at least gray) market in information.


Through Project Sentinel, my colleagues and I at Miso.ai have been mapping this market from a publisher’s perspective. We monitor more than 11,000 publisher domains and their robots.txt files to understand how they are attempting to govern AI crawlers, and whether AI crawlers even obey “do not crawl” instructions. We have run nearly one million queries through leading AI services, including Gemini, ChatGPT, Claude, Perplexity, and Microsoft Copilot, to test whether publishers’ stated crawler policies correspond to the content those services return to users.


Our latest figures remain under embargo, but the broad result is already clear: in a substantial number of cases, an AI output appeared inconsistent with the publisher’s stated policy. These are observed policy conflicts, not legal findings; the results alone don’t tell us why a conflict occurred. Nonetheless, declared policy and actual outcomes are not the same thing in this new AI economy.


Nor is the problem confined to declared AI crawlers. A new middle layer now industrializes web access for agents. Exa, Tavily, Parallel, and Ceramic provide search and retrieval APIs. Firecrawl turns websites into model-ready data. Apify provides tools to crawl sites like TheGuardian.com and Bloomberg News. Browserbase supplies cloud browsers that can navigate websites like humans. Even SerpAPI offers data collected from Google’s search engine results pages. Together, they point towards something new: a wholesale market for turning the human web into machine-ready inputs.


But they also create a new intermediary layer between knowledge producers and those consuming their work. A site may see a retrieval provider, browser session, search intermediary, or proxy address in its logs rather than the model, company, agent, or user that initiated the request.


This separation is fundamentally corrosive. Newsrooms, researchers, artists, developers, and academic publishers bear the cost of creating, verifying, maintaining, and defending their work. An intermediary can retrieve the finished product, sell access to it, and deliver a substitute without ever returning the visit, subscription, licensing payment, or direct consumer relationship that might sustain the knowledge producer. Thus, the producer bears the cost while the intermediary captures the demand. And whoever controls this demand layer and the information captured within eventually gains the power to dictate terms.


The scraping economy therefore does not merely divert revenue. It transfers market intelligence from producers to intermediaries. The intermediary is positioned to learn what users ask for, which answers satisfy them, what information they returned for, and eventually what they will pay. Meanwhile, the producer can see only a server request, if it sees anything at all. But if producers cannot see who values their work or capture any of that value, the next investigation, dataset, image, software library, documentary, or piece of expert analysis may never get made.


Policymakers around the world are waking up to this threat, starting with identity. The New York State Legislature has passed the Stealth Crawler Prohibition Act, requiring bots accessing covered news sources to disclose their operator and purpose, though the bill has not yet become law. A bipartisan Stealth Bot Prohibition Act has been introduced in the US House of Representatives, and Britain’s Automated Online Software (Access and Transparency) Bill is scheduled for a second reading in the House of Commons on October 16, 2026.


Each proposal is different, but all rest on the principle that machines should not pose as humans while extracting other people’s work covertly and without consent. These measures do not necessarily create a licensing market. But they establish a first prerequisite: machines must identify themselves before knowledge producers can grant or reject access, negotiate terms, or demand payment.


Seeking Alpha

The status quo also reflects web fracking’s dirty secret: it is wildly inefficient. Raw web pages are often verbose, redundant, stale, and poorly structured for a model’s purposes. Scraping is economically inefficient because AI systems spend tokens, compute, and engineering expertise determining which fragments matter. It is legally inefficient because crawling, training, retrieval, and other AI activities operate under a persistent cloud of potential litigation. And it is socially inefficient, because it turns publishers, researchers, nonprofits, open-source communities, public agencies, experts, artists, and writers into unwilling suppliers to a refinery they cannot inspect.


But most of all, scraping is epistemically inefficient, because most content is not “Alpha,” by which I mean information that changes the answer. In Claude Shannon’s information theory, information matters when it reduces uncertainty. In today’s information economy, Alpha is the gap between what a model already knows, what others already know, and what a particular knowledge producer can still uniquely add. Such knowledge doesn’t need to be secret or financially lucrative, but it does need to be distinct, timely, or credible enough to alter an answer, decision, or action.


A rewrite of a wire story or press release has low Alpha, whereas an exclusive interview has more. A leaked document, field investigation, original survey, newly unsealed court filing, or carefully argued op-ed is not then merely another piece of “content.” It is an uncertainty-reducing asset. Alpha matters because no competent analyst wants 12 redundant summaries of the same press release, and no serious AI research agent wants to fill its finite context window with repetitive information when the valuable part is a specific quote, data point, contradiction, revision, dissenting opinion, relationship, update, or caveat.


Agents will be insatiably hungry for Alpha at an atomic level. But to find it, they will need sources that are current, citable, attributable, rights-cleared, machine-readable, and differentiated. What then are creative industries, newsrooms, researchers, writers, and advocates of the open web to do?


The first option is to make a licensing deal, as News Corp did with OpenAI in an agreement reportedly valued at more than $250 million. But most publishers are not News Corp, and the largest AI systems do not actually need all the internet, let alone most of the news industry, to reach a sufficient level of quality. They can license a few leading sources, backfill with wire services and search crawling, and rely on the fact that most news is repeated across many different outlets anyway. The brutal market reality is that the more commoditized the information, the easier it is to arbitrage, and the less bargaining power knowledge producers retain. A deal nonetheless may be rational for a lucky few.


Another option is fortification: de-index your site, put up a hard paywall, litigate violations, and compete aggressively with your own premium data services, enterprise intelligence, and other high-value products. Some publishers, particularly in business-to-business subscription media, should do this; indeed, many already are, as they realize that a paywall that appears after an article loads is a curtain, not a lock.


But retreating behind walls and drawbridges means abandoning the core philosophy of the open web. A world where every credible source disappears behind anti-bot barricades and hard paywalls would not be a victory for journalism, research, education, culture, or democracy. It would be an open web that has given up on being read.


A third option is for knowledge producers to enter the knowledge refinery business themselves. Rather than waiting to be fracked, publishers, creators, researchers, and institutions could sell access to their own distilled knowledge systems and intelligent machine services. This “architecture of participation,” as technologist and publisher Tim O’Reilly calls it, is already emerging. The Model Context Protocol (MCP) provides an open standard for connecting AI agents to external data and services. Really Simple Licensing builds on robots.txt with machine-readable licensing and compensation terms. Cloudflare’s Pay per Crawl uses HTTP 402-style flows to make crawler access chargeable.


But none of these emerging mechanisms is a complete solution. MCP tells agents how to connect, RSL tells them what terms may apply, and payment rails make access priceable. But they cannot tell an agent what knowledge is worth buying.


Building a New Agentic Web

The first great web index asked: What does the web link to?


Google PageRank’s key insight was to ask which sites and web pages were most frequently referenced by others. By converting hyperlinks into signals of authority, it laid the foundation for a larger digital economy.


The next great web index will ask: What knowledge changes the answer?


A decentralized producer-owned Alpha index would not merely expose a cleaner copy of every article, paper, illustration, database, or creative work. Building on Reuters Institute Fellow Sannuta Raghu’s News Atom standard, it could present a machine-readable inventory of knowledge available for licensed use: the claim; the evidence supporting it; when it was verified; the confidence attached to it; its source trail; permissible uses; and the price for retrieving, quoting, summarizing, caching, or retaining it.


Humans could still read the finished article or enjoy the creative work. But the machine audience that Fang describes would receive an interface designed for evidence, provenance, and permitted use rather than the attention mechanics of a web page. And as Ilan Strauss at the AI Disclosures Project argues, open protocols are not merely technical plumbing. They shape markets: who can participate, what must be disclosed, and whether value circulates among specialized producers or remains trapped inside a handful of gatekeepers.


An Alpha index could support a genuinely two-sided information market. Agents could broadcast what they are seeking, on whose behalf they are acting, and how they intend to use the results. Publishers, researchers, creators, and other knowledge producers could respond with structured offers and knowledge graphs detailing what they know, why it is distinctive, how confident they are, what rights apply, and what price attaches to retrieval, citation, caching, or reuse.


Imagine an agent seeking evidence about which companies lobbied for changes to a national infrastructure bill. Rather than scraping hundreds of pages, it could broadcast that request. A political-intelligence publisher’s own agent might disclose that it has primary interviews, a normalized database of lobbying filings, and a verified timeline, then license underlying evidence to the agent under negotiated terms.


Such a transaction would create the very thing scraping destroys: a receipt showing who wanted the knowledge, which claims they used, what rights they received, and what it was worth. Scraping extracts what already exists, but a market can incentivize the production of new public and private goods.


Such a market, and its underlying standards and institutions, would undoubtedly disrupt today’s open web. But it would also present a rare opportunity to reassert decentralized power online. It would require an open agentic-web stack for identity, permissions, provenance, pricing, telemetry, disclosure, and machine-readable rights; much as Apache, Mozilla, the Internet Engineering Task Force, and the World Wide Web Consortium supplied the human web.


The alternative is not a neutral continuation of the status quo. AI’s appetite for knowledge and human production is not going away. Left unchecked, the open web will crumble under the weight of a handful of AI platforms, with a shadow market of information brokers extracting value underneath. The machines are already drilling. What is urgent now is to stop giving away the field. The old methods of mediating access with only humans in mind no longer work. They are too crude, and they ultimately serve as an invitation for fracking and abuse.


The open web cannot survive being fracked indefinitely. A world in which intermediaries control machine demand is one in which the very people producing our collective knowledge and art can no longer see their own customers, measure the value of what they create, or finance what comes next. Slowly but surely, the experts, journalists, researchers, developers, and artists who replenish the depths of the open web will be pushed out. And what remains may be powerfully intelligent, but hollow, detached from the human knowledge, experience, and creativity that made it possible.


Featured

Superintelligent Democracy

Aug 20, 2026Audrey Tang


When Words Are No Longer for Us

Sep 14, 2026Matthew Kirschenbaum


Who Will Do the World’s “Most Impossible Job”?

Sep 14, 2026Jorge Heine


AI Firms’ Self-Serving Warnings

Sep 15, 2026Brian Judge


Who Controls the Data Controls the Future

Sep 7, 2026Bijan Burnard


Lucky Gunasekara

Lucky Gunasekara

Writing for PS since 2026

1 Commentary


Lucky Gunasekara is CEO and co-founder of Miso.ai, an AI media lab that builds private AI systems for search, audience engagement, and research for digital publishers.


0






No comments:

Post a Comment