Perplexity Accused Of 'Stealth Crawling' Blocked Sites

Perplexity is masking its AI bots and rotating IPs to access restricted content says Cloudflare

Internet infrastructure giant Cloudflare has publicly accused AI search startup Perplexity of engaging in deceptive practices to bypass content restrictions on websites.

In a detailed blog post, Cloudflare alleged that Perplexity, which offers a conversational AI search engine, is "stealth crawling" tens of thousands of websites, even after being explicitly blocked.

According to Cloudflare, the company disguises its bots and rotates IP addresses and autonomous systems to evade detection and access content protected by anti-bot measures such as robots.txt files and Web Application Firewall (WAF) rules.

"Although Perplexity initially crawls from their declared user agent, when they are presented with a network block, they appear to obscure their crawling identity in an attempt to circumvent the website's preferences," Cloudflare's researchers wrote.

"This activity was observed across tens of thousands of domains and millions of requests per day."

To test the allegations, Cloudflare created new domains with AI bot restrictions specifically targeting Perplexity's crawlers.

The tests revealed that the startup initially identified itself using names like "PerplexityBot" or "Perplexity-User." However, when faced with a block, the crawlers reportedly changed their user agent string – the information that websites use to identify visitors – to mimic legitimate web browsers like Google Chrome on macOS.

Additionally, Cloudflare's team found that Perplexity's crawlers rotated through a variety of IP addresses not listed in its bot registry and even altered their Autonomous System Numbers (ASN) to bypass firewall rules.

The alleged tactics directly violate widely accepted internet norms, particularly the Robots Exclusion Protocol (robots.txt), a basis of web etiquette since 1994 and formalized by the Internet Engineering Task Force in 2022.

This protocol gives website operators a way to inform crawlers about what content they can and cannot access.

Cloudflare's blog post reiterates a call for clear and ethical standards when it comes to AI scraping.

"There are clear preferences that crawlers should be transparent, serve a clear purpose, perform a specific activity and, most importantly, follow website directives and preferences."

Cloudflare says it has de-listed Perplexity as a verified bot and implemented new rule sets to block what it calls "stealth crawling" across its network.

Not The First Accusation

This isn't the first time Perplexity has been accused of violating internet rules. Last year, the company was caught allegedly bypassing paywalls and ignoring robots.txt instructions to extract web content, incidents that prompted a backlash from publishers and platform operators.

At the time, CEO Aravind Srinivas downplayed the controversy, blaming unauthorized scraping on third-party bots used for testing.

Reddit CEO Steve Huffman has previously singled out Perplexity, along with Microsoft and Anthropic, for treating web content as freely available for AI training.

"We've had Microsoft, Anthropic and Perplexity act as though all of the content on the Internet is free for them to use. That's their real position," Huffman told The Verge last year.

Perplexity has dismissed Cloudflare's claims.

In a statement to The Verge, spokesperson Jesse Dwyer labeled the report a "publicity stunt" and argued that "there are a lot of misunderstandings in the blog post."

However, the company did not directly address the specific technical findings from Cloudflare's research, including the use of disguised user agents and rotating IP infrastructure.

This article originally appeared on our sister site Computing.