Close Menu
MMJ News NetworkMMJ News Network
  • Home
  • Cannabis
  • Psychedelics
  • Crypto & Web3
  • AI
  • CBD
  • Wellness & Counterculture
  • MMJNEWS

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

What's Hot

Dogecoin Down 18%, But Whale Withdraws 122 Million DOGE From Binance

September 26, 2025

This AI Chatbot Is Trained on Top Crypto Traders—Can It Offer an Edge?

September 25, 2025

Inside the Box: Aaron Levie on reinvention at Disrupt 2025

September 25, 2025
Facebook X (Twitter) Instagram
MMJ News NetworkMMJ News Network
  • Home
  • About Us
  • Advertise With Us
  • Contact Us
  • DMCA
  • Privacy Policy
  • Terms & Conditions
  • Home
  • Cannabis
  • Psychedelics
  • Crypto & Web3
  • AI
  • CBD
  • Wellness & Counterculture
  • MMJNEWS
MMJ News NetworkMMJ News Network
Home » Perplexity accused of scraping websites that explicitly blocked AI scraping
AI

Perplexity accused of scraping websites that explicitly blocked AI scraping

EditorBy EditorAugust 4, 2025No Comments3 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Share
Facebook Twitter LinkedIn Pinterest Email Copy Link


AI startup Perplexity is crawling and scraping content from websites that have explicitly indicated they don’t want to be scraped, according to internet infrastructure provider Cloudflare.

On Monday, Cloudflare published research saying it observed the AI startup ignore blocks and hide its crawling and scraping activities. The network infrastructure giant accused Perplexity of obscuring its identity when trying to scrape web pages “in an attempt to circumvent the website’s preferences,” Cloudflare’s researchers wrote.

AI products like those offered by Perplexity rely on gobbling up large amounts of data from the internet, and AI startups have long scraped text, images, and videos from the internet many times without permission to make their products work. In recent times, websites have tried to fight back by using the web standard Robots.txt file, which tells search engines and AI companies which pages can be indexed and which shouldn’t, efforts that have seen mixed results so far. 

Perplexity appears to be willingly circumventing these blocks by changing its bots “user agent,” meaning a signal that identifies a website visitor by their device and version type; as well as changing their autonomous system networks, or ASN, essentially a number that identifies large networks on the internet, according to Cloudflare. 

“This activity was observed across tens of thousands of domains and millions of requests per day. We were able to fingerprint this crawler using a combination of machine learning and network signals,” read Cloudflare’s post. 

Perplexity spokesperson Jesse Dwyer dismissed Cloudflare’s blog post as a “sales pitch,” adding in an email to TechCrunch that the screenshots in the post “show that no content was accessed.” In a follow-up email, Dwyer claimed the bot named in the Cloudflare blog “isn’t even ours.”

Cloudflare said it first noticed the behavior after its customers complained that Perplexity was crawling and scraping their sites, even after they added rules on their Robots file and for specifically blocking Perplexity’s known bots. Cloudflare said it then performed tests to check and confirmed that Perplexity was circumventing these blocks. 

Techcrunch event

San Francisco
|
October 27-29, 2025

“We observed that Perplexity uses not only their declared user-agent, but also a generic browser intended to impersonate Google Chrome on macOS when their declared crawler was blocked,” according to Cloudflare.  

The company also said that it has de-listed Perplexity’s bots from its verified list and added new techniques to block them. 

Cloudflare has recently taken a public stance against AI crawlers. Last month, Cloudflare announced the launch of a marketplace allowing website owners and publishers to charge AI scrapers who visit their sites. Cloudflare’s chief executive Matthew Prince sounded the alarm at the time, saying AI is breaking the business model of the internet, particularly publishers. Last year, Cloudflare also launched a free tool to prevent bots from scraping websites to train AI. 

This is not the first time Perplexity is accused of scraping without authorization. 

Last year, news outlets, such as Wired, alleged Perplexity was plagiarizing their content. Weeks later, Perplexity’s CEO Aravind Srinivas was unable to immediately answer when asked to provide the company’s definition of plagiarism during an interview with TechCrunch’s Devin Coldewey at the Disrupt 2024 conference.



Source link

AI Artificial Intelligence (AI) bots cloudflare LLMs Perplexity scraping
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Editor
  • Website
  • Facebook
  • Instagram

Related Posts

Inside the Box: Aaron Levie on reinvention at Disrupt 2025

September 25, 2025

Elon Musk’s xAI offers Grok to federal government for 42 cents 

September 25, 2025

Steph Curry’s VC firm just backed an AI startup that wants to fix food supply chains

September 25, 2025

Comments are closed.

Don't Miss
BTC

Dogecoin Down 18%, But Whale Withdraws 122 Million DOGE From Binance

Keshav is currently a senior writer at NewsBTC and has been attached to the website…

This AI Chatbot Is Trained on Top Crypto Traders—Can It Offer an Edge?

September 25, 2025

Inside the Box: Aaron Levie on reinvention at Disrupt 2025

September 25, 2025

Analyst Says Bitcoin Bear Market Has Started, Predicts 50% Crash To $61,000

September 25, 2025
Top Posts

Cannabis Advertising Compliance 2026: Strategies That Scale

September 25, 2025

Red Imported Fire Ants Ravage South Carolina Hemp Crop

September 24, 2025

Texas Senator Who Pushed Hemp THC Ban Now Trying to Interject in Governor’s Executive Order

September 23, 2025

Where Candidates Stand on Cannabis in Virginia, New Jersey 2025 Gubernatorial Races

September 23, 2025

Subscribe to Updates

Subscribe to our newsletter and never miss our latest news

Subscribe my Newsletter for New Posts & tips Let's stay updated!

About Us
About Us

Welcome to MMJ News Network, your premier source for cutting-edge insights into cannabis, psychedelics, crypto & Web3, wellness, counterculture, and market trends. We are dedicated to bringing you the latest news, research, and developments shaping these fast-evolving industries.

Facebook X (Twitter) Pinterest YouTube WhatsApp
Our Picks

Dogecoin Down 18%, But Whale Withdraws 122 Million DOGE From Binance

September 26, 2025

This AI Chatbot Is Trained on Top Crypto Traders—Can It Offer an Edge?

September 25, 2025

Inside the Box: Aaron Levie on reinvention at Disrupt 2025

September 25, 2025
Most Popular

Ethereum Falls as Crypto Exchange Bybit Confirms $1.4 Billion Hack

February 21, 2025

Florida Woman Accused of $850K Trump Solana Meme Coin Theft, Faces Deportation

February 21, 2025

Bitcoin, XRP and Dogecoin Sink Amid Inflation Fears and Bybit Hack Fallout

February 23, 2025
  • Home
  • About Us
  • Advertise With Us
  • Contact Us
  • DMCA
  • Privacy Policy
  • Terms & Conditions
© 2025 mmjnewsnetwork. Designed by mmjnewsnetwork.

Type above and press Enter to search. Press Esc to cancel.