Why Cloudflare Can’t Block Google From Scraping Websites For Its AI Products
Cloudflare, which powers many of the world’s most prominent websites, made waves last week by introducing a default setting for new customers that would block bots that artificial intelligence firms such as OpenAI and Anthropic use to scrape sites to train artificial intelligence.
What Cloudflare didn’t highlight was that Google’s AI products couldn’t be blocked by the new setting.
Google’s bot for collecting data for its Gemini AI models is the same one that indexes websites for Google Search, so Cloudflare can’t block it on a network level, as it is doing for other AI bots, without also cratering search traffic for website publishers. That’s a problem because many sites say they’ve already lost significant traffic and revenue due to Google’s AI products and features.
The situation shows how Google’s dominant position in search is also boosting its AI.
That could change next month when a federal judge, who ruled last year that Google runs an illegal search monopoly, will determine what changes Google needs to make. He could force Google to provide website publishers and YouTube content creators with an “easily usable mechanism” to opt out of having their content used to train any of Google’s AI products. Google is expected to appeal his ruling, no matter what it says.
Google currently offers publishers a way to opt out of having their content used to train Gemini by adding certain language to a file on their sites known as robots.txt, instructions that AI companies are expected to follow but that aren’t legally binding.
Under Cloudflare’s announcement last week, Cloudflare by default will manage robots.txt files for new customers. With the default settings, this will tell Google and other companies like Apple, which also uses the same scraping bot for search and AI products, not to use a publisher’s site for AI training.
However, there’s a difference between blocking crawlers on a network level and trusting a crawler to obey robots.txt, said Will Allen, a vice president at Cloudflare. Allen compared the former to a bouncer at a bar, checking IDs and deciding who to let in, and the latter to a “No Trespassing” sign. Startup Tollbit has found that several AI firms routinely ignore robots.txt, and firms that comply with robots.txt might still purchase website content from scraping firms that don’t comply.
Google says it complies with robots.txt and that adding language instructing it to block content scraping for Gemini wouldn’t impact how the publisher’s content appears in search results.
It’s not so simple, though. Even if publishers update their robots.txt to block Gemini, Google can still use their content to train AI-generated answers that appear at the top of search results, including AI Overviews as well as AI Mode, which lets people interact with Google Search as a chatbot that spits out answers rather than links, according to court testimony in the search monopoly case.
Some publishers say that their traffic has dropped off a cliff since Google expanded AI Overviews because it gives people less of a reason to click through to their sites. They argue Google is using their content and offering nothing in return, breaking a decadeslong practice. (The exceptions are the Associated Press and Reddit, which separately struck deals to license their content to Google to train Gemini.)
Danielle Coffey, president & CEO of the News/Media Alliance, which represents more than 2,200 news publishers including Condé Nast, The Atlantic and The Guardian, said there is “no true effective way of truly opting out” of giving Google content for its AI without completely disappearing from search results.
Cloudflare’s CEO, Matthew Prince, said on X last week that Cloudflare was working on a way to block Google from using web content for AI Overviews without impacting traditional search indexing, suggesting that the company was in conversations with Google but would also pursue a legislative approach if those talks failed. A spokesperson for Cloudflare declined to elaborate on Prince’s posts.
The situation has caused confusion among publishers. For instance, Kalee Sorey Dillard, who runs fitness and lifestyle sites with her mother, reaching around 30,000 readers per month on her main site, said she didn’t beef up the robots.txt file to block Gemini for fear of being downgraded in Google search results, despite Google’s assurances.
Tom Critchlow, executive vice president of audience growth at Raptive, which runs ad sales for thousands of web publishers and creators’ websites, said that Google’s opacity on the impacts of AI Overviews and AI Mode on publisher traffic has harmed publishers’ trust in the company.
“When Google says, you can block [its AI scraping bot] and it won’t harm your search rankings, there’s not a lot of confidence or faith that that’s actually the case,” he said.
Here’s what else is going on…
OpenAI Preps ‘Study Together’ Feature in ChatGPT
Developers are rightly waiting with bated breath for OpenAI’s launch of GPT-5 and a powerful open-weight model, but product changes in ChatGPT should get just as much attention.
Some customers have noticed a new setting called “study together” that acts like a kind of tutor or guide, asking follow up questions to tailor the information to the customer and provide examples they can learn. For instance, one customer asked ChatGPT’s study together mode to teach it about contract law and this is what it looked like.
While much has been made about students’ ability to use ChatGPT to cheat on writing or other projects for school, there’s no doubt chatbots are going to be extremely useful for both teachers and their pupils in the years to come, given the breadth of knowledge they hold.
OpenAI has seized on this opportunity, and sells a version of ChatGPT to colleges directly, with features and controls for faculty and students alike. And on Tuesday OpenAI, Microsoft and Anthropic announced they would fund a new $23 million AI education program for teachers, with hands-on training on how to use AI to make lesson plans, among other things.
I wouldn’t be surprised if one of the AI devices OpenAI is cooking up with the Jony Ive-led team it acquired can watch what a student is reading or studying at their desk and can walk them through the work.
As for the cheating thing, teachers can probably solve that by mandating a lot more live, in-person test-taking so they can find out what concepts their students actually know or retain.—Amir Efrati
HuggingFace’s Mini Robot
Machine learning company HuggingFace revealed a tiny and cheap robot on Wednesday. Standing at a formidable 11 inches tall, Reachy Mini sports two mobile antennas on its head and a simple stationary body. Though it can’t walk around, the robot can be programmed to track a user’s hand or execute dance moves.
The wired version sells for $300 and the wireless version goes for $450. That’s a steep discount from the human-size Reachy 2, which HuggingFace offers for $70,000. With Reachy Mini, the company continues to establish a niche in low-budget robots for hobbyists: it has already released $100 robotic arms that users can 3D print themselves.
Reachy was originally developed by Pollen Robotics, an open source robotics company that HuggingFace acquired in April. And fitting with HuggingFace’s penchant for open source software, users can program and share their own apps for Reachy Mini. The robot also features a camera, microphone and speaker, which HuggingFace suggests will allow AI tinkerers to use vision and speech AI models with the robots.—Rocket Drew
TI’s Deep Research
The Information on Tuesday launched its deep research chatbot, which was trained on thousands of articles to provide answers to technology and business queries, such as, “What are the biggest trends in the war between old school SaaS companies and new AI companies and who is winning.”
If you have critical feedback, we’d love to hear from you.
Deals and Debuts
See The Information’s Generative AI Database for an exclusive list of private companies and their investors.
Mistral AI is in talks to raise up to $1 billion in funding from investors including MGX, along with hundreds of millions in debt from lenders including Bpifrance, Bloomberg reported.
Boldstart Ventures raised $250 million for a new fund to invest in startups developing AI products for other companies.
SiPearl, which designs semiconductors, including for AI, raised $152.4 million in a Series A funding round from investors including EIC Fund, French Tech Souveraineté and Cathay Venture.
Arago, which is designing a light-powered chip designed to reduce the energy needs of AI chips, raised $26 million in a Series B funding round led by Earlybird, Protagonist and Visionaries Tomorrow, with participation from Generative IQ and C4 Ventures.
The American Federation of Teachers is investing $23 million to establish an academy in New York City that will provide AI training to educators, backed by funding from OpenAI, Microsoft and Anthropic.
OpenAI has revamped its security practices, such as implementing stricter controls on employees’ access to sensitive information (including through fingerprint-scanning locks) and limiting other companies’ abilities to replicate its models, The Financial Times reported.
AI video startup Moonvalley released its AI model Marey, which is trained on licensed video data.
Waymo launched teen accounts for riders aged 14 to 17 in Phoenix.
What We’re Reading
Upcoming Events
Tuesday, September 30 — AI Agenda Live NYC: The Next Wave
Save the date for The Information's next AI summit, AI Agenda Live: The Next Wave. AI Agenda Live will help attendees make sense of what’s coming in the field, from new research breakthroughs to bleeding-edge AI applications.
More detailsTuesday, October 28 — The Information’s 2025 WTF Summit
Reserve your spot for The Information’s WTF 2025 Summit. With AI reshaping business, volatile markets, and rising political uncertainty, the boldest women are coming together to lead through change.
More detailsThank you for reading the AI Agenda Newsletter! I’d love your feedback, ideas and tips: [email protected].
If you think someone else might enjoy this newsletter, please pass it forward or they can sign up here.