Reddit vs Perplexity

Just several years ago, arguments regarding scraping ended with demands to stop, written in formal letters, and IP blocks. In 2026, the situation changed: Reddit filed a lawsuit that concerns not only an AI company, but the entire data supply chain, including a proxy provider. If your team collects data, or you are looking for a data vendor or proxy provider right now, the details of that lawsuit are worth your attention. DataImpulse is a proxy provider as well, so let’s take a closer look together.

Key Facts:

  • The Reddit vs. Perplexity lawsuit is an infamous data-centered case that began in 2025 and remains ongoing in 2026. 
  • The lawsuit centers on allegations that Perplexity obtained Reddit’s human-generated data without permission and used it to train its models. Reddit says that the AI company scraped billions of Google SERPs to get data, as Reddit uses sophisticated anti-scraping measures. 
  • Perplexity denies the allegations, saying that it never trained its models on that data.
  • There are three intermediaries involved in the case: Oxylabs, AWMProxy, and SerpApi. Reddit pressed charges against them as well.
  • At the latest hearing, the court allowed Reddit to continue pressing most of its charges. 
  • The lawsuit clearly shows that now every member of the data supply chain may be held legally responsible for getting around defenses. 

Reddit vs Perplexity: What is the lawsuit about?

On October 22, 2025, Reddit, Inc. filed a lawsuit in the Southern District Court of New York against Perplexity AI, Inc. and three middlemen: Lithuania-based proxy provider Oxylabs, Russia-based network AWMProxy that relied on hacked devices, and Texas-based SerpApi, a company that provides an API for extracting raw search pages and turning them into machine-readable JSON format. According to Reddit, those companies participated in industry-scale scraping of Reddit’s human-generated content. 

The most interesting part is where the data was scraped from. Reddit claims that the data was not sourced directly from its platform. As Reddit employs a sophisticated security system against automated traffic, Perplexity apparently got the content indirectly, by scraping Google SERPs, where parts of posts resurface. The plaintiff says that there is no difference between direct and indirect data gathering if the source of data is the same protected website.

In order to prove their claims, Reddit created a trap: they published a test post that was only visible to Google’s search crawlers. That post soon appeared in Perplexity’s answers. Reddit thinks of this as evidence that Perplexity and brokers scraped Google’s SERPs in order to recreate Reddit’s posts, circumventing its access control tools. 

Another interesting fact is that there was a warning before the lawsuit. In May 2024, Reddit issued a cease and desist demanding that Perplexity stop its activities. The latter one publicly answered that the company respects policies and robots.txt files, while its Reddit citations soon increased fortyfold, according to the lawsuit. 

What is Reddit’s position?

Reddit never argues that its human-generated content is valuable for AI companies that are in a race for high-quality training data. And the platform sells it via multi-million-dollar agreements. On February 22, 2024, Reddit and Google announced their data licensing agreement, which grants Google programmatic access to Reddit’s content. Soon, OpenAI made the same deal. Reddit now provides its data for AI training. 

Reddit uses those agreements as examples of what Perplexity should have done: negotiate, sign a deal, and pay for data, instead of getting around technical defense means. Plaintiff emphasizes that Perplexity does not have a license for using Reddit’s content and partnered with at least one scraping company in order to get that data. The case itself is framed as the issue of broken terms of service and unauthorized access, not a traditional copyright infringement alone. 

After the latest hearing in the case, which took place on July 31, 2026,  a Reddit spokesperson said, “Today’s ruling ⁠brings us one step closer to holding bad actors accountable. Reddit supports responsible access to public content, but ​we oppose companies that bypass our protections, ignore our rules, and profit off our communities without permission.”

A similar lawsuit filed by Reddit against Anthropic is also ongoing as of now.

At the same time, some Reddit users voice their thoughts that the content, like comments, does not belong to Reddit either; it was generated by visitors, so is it really nice of Reddit to receive that content from nothing and sell it for millions of dollars? 

What is Perplexity’s position?

Perplexity claims that its system operates as an automated “answer engine” that responds strictly to user-initiated queries. The company also says that it does not use copyrighted data to train its models. Instead, they use retrieval-augmented generation (RAG) to pull, rank, and summarize publicly available data in real-time mode, providing users with citations and the source, and, by doing that, redirecting traffic to Reddit and popularizing it. According to the company, the whole system is acting like an advanced search engine, and there is nothing illegal about it. The company denied all allegations, and its legal team filed motions to dismiss the lawsuit.

“Reddit’s suit claims the right to control access to ​public web pages it doesn’t own, using a security tool it didn’t build, on behalf of users it hasn’t ​asked. We’re going to defend the open internet, and we’re going to win,” a Perplexity spokesperson said. 

Why are proxy providers and SerpApi also involved?

Reddit filed its lawsuit not only against Perplexity, but also against infrastructure and data providers. The platform said that Oxylabs, AWMProxy, and SerpApi collected Reddit’s data without any permission by scraping billions of SERPs. The term “data laundering” is used in the lawsuit, and middlemen are called brokers who collected and sold Reddit’s data.

AWMProxy holds a special place here. The service is described as the one tied to malware botnets like TDSS and Glupteba, linked to criminal operations and malware propagation. Cybersecurity researchers say that AWMproxy is a network of hacked computers, and its “residential” IPs belong to devices, the owners of which never gave their consent to use their connection. 

What is the current status of the case?

The key thing is that Reddit centers the lawsuit not around copying copyrighted text, but around getting around technical defense means. DMCA, the Digital Millennium Copyright Act, serves as a legal basis. It is a law that regulates copyrights in the digital era, and it contains a chapter №1201 about getting around technical defenses. Reddit claims that data companies have hidden their identities in order to bypass limitations. According to the plaintiff, two levels of protection were bypassed: Reddit’s inner antiscraping measures and Google’s control mechanisms on its SERPs. 

The latest hearing took place on July 31, 2026. Judge Paul Engelmayer has dismissed Perplexity’s and SerpApi’s motions to close the case. The court also dismissed some secondary allegations, including unjust enrichment and unfair competition charges, but advanced the rest of them. The court has agreed that GoogleSearchGuard is an access control tool according to Chapter 1201(a) of the DMCA. Among the methods Reddit claims were used is the rotation of IP addresses. 

Interesting fact: less than a couple of weeks before, another judge dismissed Google’s lawsuit against SerpApi, reasoning that Google can not demand protection of data, which it does not have the copyrights to. It shows that courts do not have a consensus yet and treat the same clauses in different ways. 

What does the Reddit vs Perplexity lawsuit mean for data teams?

One of the main tendencies is that the argument “We have only bought data from a vendor” becomes weaker and weaker. Perplexity claims that it does not work by getting around the access control measures, and the purchase of data collected by someone else is not illegal. The court, though, did not accept it. It means that there are legal risks for everyone involved in the data supply chain, from a proxy provider to the end user. 

When is proxy usage legal?

Proxies themselves are not illegal. They are a standard network tool, used by corporations, marketing agencies, security services, and QA teams. Legality is not defined by a tool; it is defined by the purposes you use it for. 

Typical legitimate scenarios include price and assortment monitoring, ad verification and dealing with ad fraud, SEO monitoring, brand protection, localized testing, and market research based on publicly available data. 

Though here the next question arises: Is it legal to collect publicly available data using proxies? Generally, access to public pages without authorization is not illegal across numerous jurisdictions. However, the Reddit vs. Perplexity lawsuit shows where the line is. Risks become drastically higher when you get over CAPTCHA or other access control tools, ignore a clear and obvious “no” from a website owner, collect personal data without legitimate interest, copy and sell copyrighted content, or violate the terms of service you accepted when creating an account. IP rotation itself is a regular practice, but if it is used to get around limitations, a court may see it in a different light.

How to choose a proxy provider?

Ethical scraping reduces legal risks drastically. Here is what you need to do. 

First, confirm the IP source. You need a provider who can prove having consent of users whose IPs are used and, preferably, pay them. 

Second, remember that an honest provider demands filling a KYC form and prohibits proxy usage for illegal scenarios. 

Third, the agreement transparency also matters. Check the division of responsibility, data processing, and GDPR compliance. 

Fourth, technical qualities like geocoverage, uptime, success rate, and speed are still important. 

Finally, pay attention to the price and payment model. Judge not only by the price per GB, but also by the price per successful request – this is a more realistic metric. Fixed subscription plans also fail often, as they do not meet the real scraping spikes, and thus make you overpay for unused traffic or buy additional GBs at a higher price. This way, you can have enterprise quality for a lower price – without risking getting involved with shadowy providers. 

To sum up 

The Reddit vs Perplexity lawsuit is still ongoing, but it has already changed the game. Platforms’ owners sell data and are ready to sue those who get around technical defenses, including intermediaries. Data teams now should document data sources, respect robots.txt and limitations, react to website owners’ complaints, and work only with trusted third parties that can prove the legal origin of their IPs or data. 

This article is written for informational purposes only and does not replace a legal consultation. 

Frequently Asked Questions

What is the Reddit vs Perplexity lawsuit about?

The Reddit vs Perplexity lawsuit is centered around allegations that Perplexity, with the help of three intermediaries, scraped Reddit's data without permission and used the data for training its model. According to Reddit, data was scraped from Google SERPs, bypassing both Reddit's and Google's defenses.

How did the Reddit vs Perplexity lawsuit end?

The lawsuit is still ongoing. The latest trial took place on July 31, 2026, and the court allowed Reddit to continue pressing most of its charges.

What intermediaries are involved in the Reddit vs Perplexity lawsuit?

There are three intermediaries, and there is a proxy provider among them. There is also an API provider, SerpApi, which offers an API to extract raw data and turn it into JSON format.

Is it legal to use proxies?

The Reddit vs Perplexity lawsuit is not about proxy usage, but about circumventing limitations set by a platform. Proxy usage itself is legal as long as you do not try to get around access control means or harvest data in a way that violates the terms of service of the target website.

What is DataImpulse?

DataImpulse (dataimpulse.com) is an ethical vendor of residential, datacenter, mobile, and premium residential proxies. The vendor offers 90M+ residential, 20M+ datacenter, and 16M+ mobile IPs from 195 countries on a pay-per-GB model, and its traffic is non-expiring. The human support team is on standby 24/7. The main use cases include price monitoring, geo testing, ad verification, and SERP tracking.

Is DataImpulse legal?

Yes, DataImpulse is a legal provider. It provides 1st party residential IPs, sourced with users' consent. Its proxies are ISO certified and GDPR compliant. The provider uses 2FA, KYC, and prohibits the use of its IPs for illegal purposes, so your reputation stays safe.

When is DataImpulse not the best choice?

DataImpulse does not support the use of its proxies for any illegal activities. Also, the vendor does not provide access to government or banking sources, a ready-to-use scraping API, or static addresses.

Share article: