Is Perplexity citing spam farms on purpose?
SUMMARY
No. The evidence points to Perplexity repeatedly letting spam-like publishers into its retrieval layer, not deliberately promoting them because it wants low-quality sources in its answers.
The scale is still uncomfortable. Four unusual high-volume publishers generated about 5% of all citations in one 7,534-citation software audit, and Guideflow alone appeared across roughly one quarter of the software categories tested.
The clearest spam-farm examples look purpose-built for machine retrieval. WiFiTalents, WorldMetrics and Gitnux collectively created more than 215,000 “best software” URLs, used closely related templates and infrastructure, and even described some pages with machine-oriented language such as “Facts & Grounding Page.”
Perplexity’s problem is especially visible on obscure commercial queries. When someone asks about elections, established sources are plentiful. Ask for the best marina-management CRM or arborist software and the evidence pool suddenly fills with vendors, affiliates, directories and mass-produced rankings.
Perplexity also reaches unusually deep into the web. In one 5,000-query dataset, only 39.4% of its cited URLs appeared in Google’s organic top ten for the same query. That ability to find obscure specialists is useful, but it gives manufactured publishers another way into answers.
The bigger concern may actually be citation provenance. A separate Haus Research audit found that 34.7% of 1,826 numerical citation markers either could not be retrieved or pointed to readable pages containing none of the figures in the sentence they supposedly supported.
There is still an important missing experiment. Current research shows that spammy pages enter Perplexity’s evidence layer, but nobody has yet removed those pages and rerun the same software questions to measure whether the recommended products actually change.
Perplexity has known about this class of problem since at least 2024. The company was already discussing trust scores, spam downranking and AI-content detection then, so repeated contamination today is hard to describe as a completely unexpected failure.
The “on purpose” theory gets weaker when money enters the discussion. Perplexity does have publisher revenue-sharing arrangements, but there is no evidence that its ranking system deliberately favors unpaid spam publishers to avoid sharing advertising revenue.
The more convincing criticism sits one level higher in the product design. Perplexity deliberately favors deep retrieval, broad source coverage, permissive access to unrated domains and lots of visible citations. Those choices make the product useful, and they also create an unusually large surface for publishers trying to game AI search.
So Perplexity probably does not want spam farms in its answers. But after years of knowing that synthetic and low-quality pages can enter its citations, calling the current pattern a surprising accident is getting difficult too.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →What exactly is Perplexity citing that looks like spam?
Perplexity is currently citing a small but surprisingly visible group of mass-produced software-ranking sites alongside normal company pages, review sites and established publishers.
A Trellner Research audit published this week ran 380 software-buying questions through Perplexity’s Sonar and Sonar Pro models. Across 760 answers, the models produced 7,534 citations from 2,055 different domains. Almost 60% of those citations came from sites outside the 100,000 most visited domains in Tranco, and 23.4% came from domains outside its top million.
Low traffic alone tells us very little about quality. Plenty of tiny websites contain excellent specialist information. The strange part appears when we look at which obscure domains kept coming back.
Guideflow, a real company selling interactive product demos, received 194 citations across 96 of the 380 software categories. Then Trellner found another 181 citations pointing to WiFiTalents, WorldMetrics and Gitnux. Those three sites had collectively published 215,128 URLs following a /best/...-software/ pattern.
That is what triggered the current controversy. Perplexity occasionally stumbling onto a weird blog would barely be interesting. A handful of aggressively manufactured content estates repeatedly entering the evidence behind hundreds of software recommendations is much harder to dismiss.
Are those Perplexity sources actually spam farms?
The strongest examples in Perplexity look much closer to industrial content farms than ordinary niche blogs.
We should separate Guideflow from the stranger sites. Guideflow is a real software company with employees, customers and a product. Its huge collection of comparison pages looks like aggressive content marketing. Calling the entire company a spam farm would stretch the term too far.
WiFiTalents, WorldMetrics and Gitnux are different.
Trellner found that the three domains used closely related templates, shared infrastructure, cross-linked between themselves and appeared within a relatively narrow registration period. Each had roughly 70,000 software-ranking pages. Some pages described themselves in their HTML metadata as machine-readable records, while WorldMetrics and Gitnux used the title “Facts & Grounding Page.”
The content itself makes the pattern even stranger. Equivalent pages across the network could recommend completely different products while still presenting the rankings as “AI-verified” or “expert reviewed.” Trellner also found unfinished template text appearing publicly on some pages.
A normal editorial operation can publish thousands of useful pages. Publishing roughly 70,000 near-identical “best software” pages per site, with explicit machine-oriented wording and templated rankings across obscure categories, points toward content manufactured for retrieval at scale.
That is close enough to what most people mean when they say “spam farm.”
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →How big is the spam-source problem in Perplexity?
Perplexity’s documented spam-farm problem is large enough to affect its evidence layer, while the clearly identified farms still represent a minority of all citations.
Guideflow plus WiFiTalents, WorldMetrics and Gitnux generated 375 citations in Trellner’s dataset. That works out to almost exactly 5% of all 7,534 citations.
Five percent sounds modest until we look at distribution. Guideflow appeared in 96 of 380 software categories, meaning roughly one in four categories touched the same marketing publisher. The three closely linked “best software” sites appeared far less often, yet they still managed to enter dozens of Perplexity answers.
The larger long tail deserves attention too. Trellner checked cited domains against both Tranco and the Internet Archive. Among cited domains outside Tranco’s top million that had archival records, 16.6% had first appeared in the Internet Archive in 2025 or later. Among ranked domains, only 1.6% were that young.
So the obscure websites entering Perplexity are disproportionately new. That gives us no right to call every young domain spam, but the tenfold gap makes the source pool unusual enough to investigate seriously.
| Perplexity sourcing measure | Result | What it tells us |
|---|---|---|
| Citations reviewed | 7,534 | Large enough to see recurring patterns |
| Citations outside Tranco top 100,000 | 59.8% | Perplexity reaches deep into the long tail |
| Citations outside Tranco top 1M | 23.4% | Very obscure sites form a meaningful share |
| Four unusual high-volume publishers | 375 citations, about 5% | A few content estates can gain real visibility |
| Guideflow category coverage | 96 of 380 | One marketing publisher reached about 25% of categories |
| Very young unranked vs ranked domains | 16.6% vs 1.6% | The obscure source pool is much younger |
Is Perplexity worse than ChatGPT or Gemini at citing AI-written pages?
Perplexity currently looks better than Copilot and Gemini on AI-written sources, but slightly worse than ChatGPT in one of the largest cross-engine audits we have.
Northwestern University researchers Mowafak Allaham and Nicholas Diakopoulos tested 712 real questions covering health, politics and the environment. They collected more than 26,000 cited URLs from ChatGPT, Copilot, Gemini and Perplexity, then analyzed successfully retrieved pages for evidence of AI-generated writing.
Across the four systems, around 16% of the analyzed sources were classified as AI-generated.
Perplexity came in at 9.4%. ChatGPT was lower at 7.3%, while Gemini reached 14.7% and Copilot 27.8%.
The researchers also tested the detector against articles already documented as AI-generated and found a substantial false-negative rate. Their percentages should therefore be read as a floor rather than a perfect count of synthetic content.
This result pushes against the easy version of the story. Perplexity clearly has an AI-source problem. The available evidence gives us little reason to believe it has the worst one.
The much uglier results from the software audit probably have a lot to do with what people are searching for.
| AI search engine | Sources classified as AI-generated |
|---|---|
| ChatGPT | 7.3% |
| Perplexity | 9.4% |
| Gemini | 14.7% |
| Copilot | 27.8% |
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Why does Perplexity keep finding obscure websites?
Perplexity keeps finding obscure websites because its search system reaches much further down the web than the handful of sources already dominating Google.
A separate 2026 dataset gives us a useful sense of scale. Researcher Danish Shashko ran 5,000 queries across six AI-search platforms and collected 153,425 citations. Perplexity accounted for 8,562 of them.
Only 39.4% of Perplexity’s cited URLs also appeared in Google’s organic top ten for the same query.
That leaves roughly six in ten citations coming from somewhere else.
Interestingly, Perplexity still behaved more like conventional search than most competing AI systems in that dataset. At the domain level, just over half of its cited sites appeared somewhere in Google’s top ten results. ChatGPT’s URL overlap was only 4.2%.
Perplexity therefore seems to combine a conventional search backbone with much deeper retrieval. That can be excellent when someone wants an obscure technical answer hidden on a specialist website. It also gives publishers another route into the answer without first becoming a famous or heavily linked domain.
Now imagine two pages answering “best marina management software.”
One comes from a tiny consultant who has actually used ten products. The other comes from a site that automatically created 70,000 pages covering every possible software category.
A retrieval system still has to tell those two apart.
That is the hard part Perplexity has yet to solve reliably.
Are websites now building pages specifically for Perplexity and AI search?
Yes, some publishers are clearly building pages for machine retrieval, and Perplexity is one of the systems those pages can reach.
The clearest evidence comes directly from the pages themselves.
WorldMetrics and Gitnux used “Facts & Grounding Page” in page titles. Their metadata described pages as machine-readable. The wider three-site network produced huge numbers of pages using rigid structures around predictable commercial questions.
As seen above, the network reached 215,128 “best software” URLs. That scale makes far more sense when every page is treated as a potential entry point into search engines, LLM retrieval systems and recommendation answers.
The important distinction concerns targeting.
We have evidence that these publishers want machines to find and reuse their pages. We have no evidence that they built the sites exclusively for Perplexity.
The same page can potentially surface through Google AI Mode, ChatGPT, Gemini, Copilot, Perplexity and dozens of smaller products that use web retrieval. A whole industry now talks openly about GEO, AEO and “LLM optimization.” Publishers are learning how to make sentences easy for answer engines to extract, how to structure comparison pages and how to cover thousands of narrow prompts.
Perplexity happens to make the result especially visible because it shows the sources.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Did Perplexity already know spammy sources were getting through?
Perplexity has known for more than two years that AI-written and spammy pages were slipping into its citations.
Forbes and GPTZero documented the problem back in June 2024. GPTZero tested more than 100 prompts and reported that users needed only around three searches on average before encountering a source it classified as AI-generated.
The examples were broad. Travel questions surfaced AI-written blogs. Technology questions pulled synthetic material. One query asking about alternatives to penicillin cited an apparently AI-generated clinic article containing contradictory medical information.
Perplexity did not brush the underlying issue aside.
Its chief business officer Dmitry Shevelenko told Forbes that the system was “not flawless.” He said Perplexity assigned trust scores to sources, downranked websites containing large amounts of spam and had developed internal systems to detect AI-generated content.
That response is important now.
A new problem can reasonably catch a search company by surprise. Two years later, repeated contamination by machine-made sources looks more like an unresolved weakness in a system the company already understands.
The fresh Trellner and Haus Research audits make “Perplexity simply did not know this could happen” very difficult to defend today.
Are Perplexity’s new source labels actually stopping weak sources?
Perplexity’s new source labels make strong domains easier to recognize, but weak unlabeled pages can still enter answers.
Perplexity recently introduced visible labels including Government, Academic and Trusted. Reuters can receive a Trusted label, for example, while scientific publishers can appear as Academic and official agencies as Government.
The review looks at things such as named authors, correction practices and whether a publication clearly separates reporting from advertising and opinion.
This sounds like exactly the kind of layer that could solve the spam problem. The catch sits in Perplexity’s own explanation of the feature: most websites currently have no label, and an unlabeled site remains eligible to appear.
The labels also apply to domains rather than individual claims. A generally reliable publication can still publish a weak page. An obscure but excellent specialist can remain unrated. Perplexity therefore uses the labels as information for users rather than a strict citation whitelist.
That design choice is understandable. A system that only cited pre-approved domains would become terrible at obscure questions very quickly.
It also means the source labels cannot keep a newly created 50,000-page content operation outside the retrieval layer by themselves.
Perplexity says partnerships and payments do not influence these labels, which removes one obvious commercial conflict. The remaining challenge is technical: deciding whether an unrated page deserves to be treated as evidence before a human has ever reviewed its domain.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Are publisher blocks making Perplexity’s source pool worse?
Publisher blocking probably makes Perplexity’s source pool noisier, particularly in areas where reputable sites aggressively restrict AI crawlers.
A 2025 study looked at how reputable websites and misinformation sites configure robots.txt for AI crawlers. The difference was enormous.
Around 60% of reputable sites blocked at least one AI crawler. Only 9.1% of misinformation sites did the same. Reputable sites blocked an average of 15.5 AI user agents, while misinformation sites blocked fewer than one.
The gap had also widened sharply over time.
Perplexity currently says its PerplexityBot indexing crawler respects robots.txt. Its developer documentation separately describes a Perplexity-User agent used when a person asks Perplexity to visit a webpage, and says those user-requested fetches can operate differently from ordinary crawling.
Whichever access route we look at, the structural problem remains.
High-quality publishers have strong reasons to restrict automated access: copyright, licensing, subscriptions and control over how their journalism gets reused. A content farm hoping to appear in AI answers has the opposite incentive. It wants every crawler through the door.
That creates a distorted web for retrieval systems.
Perplexity still owns the ranking decision, so crawler availability cannot excuse weak citations. But the company increasingly has to rank sources inside a pool where some of the best publishers are deliberately making themselves harder to retrieve and the most aggressive machine-oriented publishers are making themselves extremely easy to retrieve.
Do Perplexity citations actually prove the claims beside them?
Perplexity’s bigger sourcing problem today may be provenance: a surprisingly large share of displayed citations fail to prove the sentence beside them.
Haus Research published a particularly useful audit this week because it avoided fuzzy judgments about whether a site “looks trustworthy.”
The researchers asked Sonar and Sonar Pro 310 factual questions about 210 technology companies. They collected citations attached to claims about funding, revenue, headquarters, CEOs, entry prices, headcount, acquisitions and other facts.
For claims containing numbers, they used a deliberately forgiving test: does the cited page contain even one of the figures stated in the sentence?
They checked 1,826 citation markers.
In 34.7% of cases, the page either could not be opened by their retrieval process or opened without containing a single figure from the sentence it supposedly supported.
Dead links explained very little of that result. Only 1.3% were genuinely dead. The larger problems came from gated or inaccessible pages and readable pages whose content simply failed the numerical test.
Haus also found that 23.1% of Sonar’s citations pointed to B2B directories, revenue estimators or lead-list sites such as Tracxn, PitchBook, Clay, GetLatka, ZoomInfo, CB Insights and similar services. Another 23.4% pointed directly to the company being discussed.
Measured at the claim level, the result looked better because Perplexity often attaches several sources to the same sentence. A claim counted as passing if at least one citation contained one of its figures. Even then, 14.4% of the numerical claims failed.
| Haus Research test | Result |
|---|---|
| Numeric citation markers checked | 1,826 |
| Citation markers failing the basic provenance test | 34.7% |
| Numeric claims failing when any supporting source could rescue them | 14.4% |
| Dead cited URLs | 1.3% |
| Sonar citations from B2B directories / lead lists | 23.1% |
| Sonar citations from companies’ own domains | 23.4% |
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Are spam farms changing Perplexity’s software recommendations?
We still cannot show that spam farms are changing Perplexity’s final software rankings, and this is the biggest missing experiment in the current evidence.
Trellner measured which sources entered Perplexity’s evidence layer. The researchers never blocked WiFiTalents, WorldMetrics, Gitnux or Guideflow and then reran the same questions.
That counterfactual would tell us much more.
Imagine Perplexity recommends products A, B and C while citing one spammy ranking page. Remove that source and the answer still says A, B and C. The citation was ugly, but it had little measurable influence on the recommendation.
Now imagine removing the page makes product C disappear and product D enter instead. Suddenly the content farm has demonstrated causal influence over the shortlist.
We do not have that experiment yet.
Trellner did find something else worth knowing. Sonar and Sonar Pro returned byte-for-byte identical citation lists in 289 of the 380 software categories, and their cited URL sets had a Jaccard overlap of 0.898. Both models also chose the same number-one product in 290 categories.
That suggests the two products rely heavily on a shared retrieval layer. When a manufactured source enters that layer, switching from Sonar to Sonar Pro may do relatively little to escape it.
So the current evidence proves source contamination much more strongly than recommendation manipulation.
Anyone claiming that spam farms are already controlling which software Perplexity tells people to buy is running ahead of the data.
Is this mostly a problem with commercial software searches?
Perplexity’s ugliest spam-farm exposure currently appears on commercial long-tail searches, especially obscure software categories.
That helps explain why different audits can make Perplexity look like two different products.
Trellner deliberately tested questions such as the best software for hundreds of specific business needs. These markets are perfect for mass-produced content. Many categories have no obvious authoritative source, while vendors, affiliates, directories, consultants and SEO sites all publish rankings.
Compare that with politics.
During the 2024 UK general election, the Reuters Institute tested Perplexity and ChatGPT on election questions. Perplexity’s answers were judged correct 83% of the time. Its citations frequently came from established news organizations, public authorities and fact-checkers. The researchers also found that Perplexity usually linked directly to pages relevant to the claims being made.
A broader 2025 study of more than 366,000 citations from OpenAI, Perplexity and Google systems found low-credibility news sources were rarely cited.
These results fit together once we stop treating “the web” as one source market.
Ask who won an election and Perplexity can choose between Reuters, the BBC, government websites and official election data.
Ask for the best arborist scheduling platform or marina-management CRM and the source universe gets messy very quickly.
Spam farms thrive in that second environment because authoritative supply is thin while commercial content supply is almost unlimited.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Does Perplexity have a business reason to cite spam farms?
We found no evidence that Perplexity makes more money by deliberately choosing spam farms.
There is a tempting theory here.
Perplexity has publisher agreements under which participating publications can receive a share of advertising revenue when their work is cited in an answer where an ad appears. Digiday reported that participating publishers can receive a percentage of that revenue, with the amount depending on their agreement.
An unpaid content farm presumably receives nothing.
From there, someone could jump to the claim that citing cheap or unpaid sources saves Perplexity money.
The evidence does not get us anywhere near that conclusion.
We found no ranking document, leaked instruction, employee statement, experiment or statistical pattern showing that Perplexity boosts unpaid publishers to reduce revenue-sharing costs. Its current source-label policy explicitly says partnerships, payments and other commercial relationships do not affect whether a domain gets a Government, Academic or Trusted label.
That statement only covers source labels, so it cannot prove that every part of ranking is commercially neutral. Still, accusing Perplexity of deliberately selecting spam to avoid publisher payments would require evidence we simply do not have.
The financial motive is currently the weakest part of the “on purpose” theory.
Is Perplexity deliberately choosing more citations even when quality drops?
Perplexity appears to be deliberately optimizing for broad sourcing, and that choice creates a real quality tradeoff.
Haus Research gives us a neat comparison.
Sonar and Sonar Pro each averaged about 9.8 cited sources per answer in its technology-company test. A GPT-4.1 web-search control averaged roughly two.
Almost five times as many sources sounds reassuring at first.
Sometimes it genuinely is. If several independent sources all confirm a funding round or product price, the answer becomes easier to trust.
The trouble starts with the eighth, ninth and tenth source when the best evidence has already been used.
Haus found almost a quarter of Sonar’s citations coming from B2B directories, estimators and lead-generation databases. The latest Trellner audit independently found mass-produced software pages reaching the citation layer. Put those findings together and a fairly intuitive pattern appears: Perplexity has a strong appetite for sources, while the long tail available to satisfy that appetite gets progressively messier.
Nobody at Perplexity has to sit down and decide, “Let’s cite bad websites,” for that outcome to happen.
The intentional decision happens one level earlier. Perplexity wants broad retrieval, lots of grounding and visible citations. Those features are core to what differentiates the product.
The quality cost shows up downstream when retrieval keeps going after the strongest evidence has already been found.
More citations are useful only while the marginal citation adds information worth trusting.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Is spam in citations really a Perplexity problem or an AI-search problem?
Spammy and AI-written citations are an AI-search problem across the board, although Perplexity deserves extra scrutiny because sourcing sits at the center of its product.
The Northwestern audit found AI-generated sources inside every system it tested: ChatGPT, Perplexity, Gemini and Copilot.
The robots.txt research gives us another industry-wide mechanism. Reputable sites increasingly restrict AI crawling while misinformation sites remain much more open. At the same time, publishing AI-generated pages has become almost free.
Those two forces are moving in opposite directions.
The expensive web is becoming harder to access. The cheap web is exploding in volume.
Every answer engine has to rank inside that environment.
Perplexity’s exposure feels more obvious because its citations are constantly visible. That visibility is actually useful for researchers. Trellner can collect 7,534 Perplexity citations. Haus can connect a sentence directly to the page supposedly proving it. We can trace exactly which domains keep entering answers.
A chatbot that absorbs equally weak material without showing a source trail could be harder to catch.
Still, Perplexity has built much of its reputation around the idea that citations make AI answers more verifiable. CEO Aravind Srinivas once described citations as the company’s “currency.”
That raises the standard.
When the citation itself becomes unreliable, the problem hits the feature that is supposed to make Perplexity safer than an ordinary chatbot.
So is Perplexity citing spam farms on purpose?
No. The evidence points to a known retrieval weakness rather than a deliberate policy of promoting spam.
We found nothing showing that Perplexity identifies a domain as a spam farm and then intentionally boosts it. There is no demonstrated payment relationship with the sites Trellner uncovered, no leaked ranking rule rewarding spam, and no convincing financial motive for preferring those sources.
Calling the current pattern an innocent one-off would also be far too generous.
Perplexity was publicly discussing AI-generated and spammy citations back in 2024. The company said it already had trust scores, spam downranking and AI-content detection. It has since added visible source labels and clearer source-quality controls.
Yet the problem remains easy to measure today.
Fresh audits show machine-oriented software sites repeatedly entering Perplexity’s retrieval layer. Another fresh audit finds serious gaps between claims and the citations displayed beside them. Independent research shows that AI systems are working inside a web where reputable publishers increasingly restrict crawlers while synthetic content gets cheaper and easier to publish.
Perplexity also deliberately keeps several design choices that make this attack surface larger: deep web retrieval, broad source coverage, permissive access to unrated domains and a preference for grounding answers with many citations.
Those choices make Perplexity useful. They also make it gameable.
So the sharpest answer to “Is Perplexity citing spam farms on purpose?” is mostly no, if “on purpose” means Perplexity actually wants spam in its answers.
The criticism becomes much stronger when “on purpose” refers to the system design. Perplexity knowingly runs a broad retrieval architecture that spam publishers can exploit, and the company has known about that weakness for years.
Calling the spam citations deliberate promotion goes too far.
Calling them a surprising accident now goes too easy.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →OUR METHODOLOGY
This analysis tests whether Perplexity is deliberately citing spam farms or whether the pattern is better explained by weaknesses in how its retrieval and citation systems work. We looked separately at how often unusual sources enter Perplexity’s answers, what those sources look like, how Perplexity compares with other AI search systems, whether displayed citations support nearby claims, and whether there is evidence that weak sources actually change recommendations.
We treated three claims separately: a questionable source can enter Perplexity’s retrieval layer; that source can materially influence an answer; and Perplexity can deliberately favor that source. Citation audits can establish the first. Showing influence properly requires counterfactual testing, such as removing a source and rerunning the same recommendation. Deliberate promotion would require stronger evidence such as ranking rules, internal instructions, demonstrated financial incentives or systematic preferential treatment.
We prioritized recent empirical audits, original datasets, reproducible research and Perplexity’s own documentation. Older reporting was mainly used to establish when Perplexity first publicly acknowledged problems with AI-generated or spammy citations. Traffic rank, domain age and AI-content detection were treated as supporting indicators rather than proof of spam by themselves.
For the software-specific evidence, we relied heavily on Trellner Research’s audit of 380 software-buying questions and Haus Research’s citation-provenance audit. We also used cross-platform research to compare Perplexity with ChatGPT, Gemini and Copilot, and separate research on AI-crawler blocking to understand how the available web itself may be changing.
We did not assume that a suspicious citation changed a recommendation. That remains one of the largest gaps in the current evidence. Likewise, Perplexity’s publisher revenue-sharing arrangements were examined as a possible incentive, but we found no evidence connecting those arrangements to deliberate ranking of spam publishers.
Key sources used for this analysis include: Trellner Research’s audit of manufactured sources in Perplexity software recommendations, Haus Research’s Sonar and Sonar Pro citation-provenance audit, the Northwestern cross-engine study of AI-generated cited sources, GPTZero’s investigation of AI-generated sources appearing in Perplexity, Forbes on Perplexity’s 2024 response to spammy AI sources, the 153,425-citation cross-platform dataset, OrganiKPI’s analysis of that citation dataset, Tranco’s popularity-ranking methodology, Perplexity’s documentation on source labels, Perplexity’s crawler documentation, Perplexity’s explanation of how its search works, research on AI-crawler blocking by reputable and misinformation sites, the Reuters Institute study of AI chatbots during the 2024 UK general election, the large study of news-source citation quality across AI search systems, and Digiday’s reporting on Perplexity’s publisher revenue-sharing system.
Get the biggest database of
profitable internet businesses
We mapped 300+ proven digital businesses so you can skip the blind trial and error. For each one, you get the site, the revenue numbers, the distribution strategy, the repeatable patterns, and ideas to recreate the model in a different niche, channel, or angle.
Get the full database →Related blog posts
- Should you write listicles about your business?
- Should you do YouTube SEO for your SaaS?
- Why is Outbid getting labeled as spam by Ahrefs?
- Do startup directories still work for SEO?
Who wrote this?
STEAL WHAT WORKS TEAM
We study profitable internet businesses, take them apart, and write down what actually works: pricing, distribution, growth, packaging. We turn 300+ proven examples into a database so founders can stop testing random ideas and start from proof. Explore the database →