AI Search Citations vs Recommendations: The Gap Your SEO Team Is Not Measuring
AI search citations and AI recommendations are two different lists, and almost nobody measures the second one. I ran 12 commercial buying questions through five real AI answer engines in a logged-in browser, then ran the whole set again. For every answer I recorded two separate things: which domains it cited, and which brands it named in the text.
Those two lists barely overlap.
Across 46 answers the engines cited 121 distinct domains and named 28 brands. 98 of the 121 cited domains — 81% — were never named once. Meanwhile Serper.dev was named in 35 answers and cited in 7.
If you are optimising for AI search on the assumption that getting cited is the goal, this is the number that should worry you. Citation and recommendation are two different games, and most of the industry is measuring the wrong one.
TL;DR: AI answer engines cite one set of sources and recommend another. Across 46 real AI answers covering 12 SERP-API buying questions, 121 domains were cited and 28 brands were named; 81% of cited domains were never named. The most-cited domain in the whole sample was named in only 8 of the 46 answers. A brand cited in just 7 answers was named in 35. Being retrieved does not make you the answer. The full dataset is public.
Two metrics, not one
Almost every AI visibility tool reports one number: were you cited. That collapses two very different outcomes into one score.
Cited means a link to your domain appeared in the answer's source list or as an inline chip. The engine read you.
Named means your brand appears in the sentence the reader actually acts on — "use X for this" — with every anchor and citation chip stripped out first.
That stripping step is the whole ballgame, and it is where most measurement goes wrong. Citation chips contain brand names. If you grep the rendered answer for your brand, the chips match and you score yourself as recommended when the answer never mentioned you. On one Claude answer in this study a naive brand grep returned four hits — all four were citation chips, and the prose named seven competitors instead.
So: strip the anchors first, then count. The two numbers that fall out are what this index is built from. This is a different cut from citations versus classic rankings, and a different cut again from AI Overviews versus AI Mode source overlap — both of those compare citation lists to each other. This one compares the citation list to the recommendation.
How I measured AI search citations and recommendations
Disclosure, up front: I run Serpent API and apiserpent.com is in this index. I did not exclude it, and the result is not flattering: it is the most-cited domain in the sample and one of the least-recommended. Every number below is reproducible from the published dataset. Judge the method, not the messenger.
- Queries: 12 commercial questions a developer actually types when choosing a SERP API — best serp api, cheapest serp api that is still reliable, serpapi alternatives, best web search api for ai agents, best google shopping api and eight more.
- Surfaces: Google AI Mode, Perplexity, Claude, ChatGPT and Google Gemini — the consumer products, in a real logged-in browser, not a vendor API.
- Answers analysed: 46 for the citation index, across four surfaces. Gemini is excluded from it entirely. The index is a joint measurement — citation and naming observed on the same answer — and Gemini exposes no source list at all, so one of the two can never be observed there. A surface where half the measurement is structurally impossible cannot join the sample.
- ChatGPT ran in temporary chat. This matters more than anything else in the setup. In a normal chat, ChatGPT referenced my earlier sessions verbatim — "since you've been comparing cheap SERP APIs" — and put my own brand first. Re-run clean, it did not mention it at all. Any AI visibility number captured in a personalised chat is measuring your own history.
- Cited = every outbound link in the rendered answer, deduplicated to the link host with
www.stripped. Subdomains are counted separately, sodocs.example.comandexample.comare two entries — the code below does exactly this. Named = brand strings in the answer body after removing every anchor and chip. - Date: captured 31 August 2026, single geography, one browser profile. Re-check before quoting; AI answers move.
The AI Citation Index: cited vs named
Here is the top of the index. Cited is how many of the 46 answers linked to that domain. Named is how many named that brand in the answer body. A domain appears here if it cleared three of either.
| Domain | Cited (of 46) | Named (of 46) | Gap |
|---|---|---|---|
| serpapi.com | 20 | 40 | +20 |
| dataforseo.com | 17 | 37 | +20 |
| apiserpent.com (ours) | 30 | 9 | −21 |
| serper.dev | 3 | 34 | +31 |
| brightdata.com | 10 | 26 | +16 |
| scrapingdog.com | 15 | 15 | 0 |
| cloro.dev | 16 | 5 | −11 |
| oxylabs.io | 0 | 18 | +18 |
| searchapi.io | 4 | 13 | +9 |
| scraperapi.com | 4 | 12 | +8 |
| openwebninja.com | 11 | 2 | −9 |
| scrapebadger.com | 11 | 0 | −11 |
Read the gap column, not the citation column. Serper.dev is the extreme case: cited in 7 answers, named in 35. It is recommended five times more often than it is read. Scrapebadger is the mirror image — cited in 6 answers, named in none.
81% of AI search citations never become recommendations
Zoom out from the leaderboard and the pattern is not a quirk of a few brands. It is the shape of the whole dataset.
Across the 46 answers, 121 distinct domains were cited. 28 distinct brands were named. Only 23 domains managed both. 98 cited domains — 81% — were never named in a single answer.
That long tail is doing real work. It is comparison blogs, Reddit threads, vendor docs, cost calculators, YouTube videos and GitHub repos. The engines read all of it to form an opinion. Then they hand the credit to a short list of brands they already knew.
Named without ever being cited
The most direct evidence that recommendations come from memory rather than from reading: five brands were named in answers that never cited them once, and the wider pattern is bigger than that handful: the brands with the largest gaps are cited a little and named a lot.
| Brand | Answers naming it | Answers citing it |
|---|---|---|
| ValueSERP | 4 | 0 |
| Decodo | 3 | 0 |
| Smartproxy | 1 | 0 |
| HasData | 1 | 0 |
| ZenRows | 1 | 0 |
The sharper version of the same point is Serper.dev: recommended in 35 of 46 answers while being linked as a source in 7, and Oxylabs, recommended in 11 while linked in 1. Whatever produced those recommendations, it was not the pages the engine had just read.
This is the practical ceiling on content-led AI visibility work, and nobody selling an AI visibility dashboard puts it on the box. You can earn your way into the source list with better pages. Earning your way into the sentence is a slower, different problem — it is brand memory, and it is measured in years of being the obvious answer.
The per-engine split
| Engine | Answers | Median sources | Brands named per answer | SerpApi cited / named |
|---|---|---|---|---|
| Google AI Mode | 12 | 8 | 5.7 | 8 / 11 |
| Perplexity | 10 | 10 | 5.9 | 3 / 10 |
| Claude | 12 | 4.5 | 6.7 | 5 / 11 |
| ChatGPT | 12 | 3 | 5.2 | 6 / 12 |
Two things stand out.
ChatGPT links to almost nothing. A median of 3 source domains per answer, against 8 for Google AI Mode and 10 for Perplexity — while still naming about as many brands. It is the surface where the citation-to-recommendation link is weakest, and the one where being cited helps you least.
ChatGPT named SerpApi in all 12 of its answers while linking to serpapi.com in 6, and Claude named it in 11 of 12 while linking to it in 5. That single row is the study in miniature: the recommendation is stable across every question, the citation is incidental, and the two have almost nothing to do with each other. It also lines up with the citation-rate gap between Claude and AI Overviews measured earlier this year.
The awkward part: we are in our own index
I would rather report this than have someone else find it.
apiserpent.com was cited in 26 of 46 answers — more than any other domain in the study, including SerpApi's 22 — and named in 8. On 73% of the answers that pulled our pages, the answer went on to recommend somebody else.
Some individual cases are stark. On best serp api free trial, Claude linked to our pages 12 times, ranked us second of its 14 sources, and then recommended SerpApi, DataForSEO, Serper.dev, Bright Data, Scrapingdog, ScraperAPI and Tavily — every one of them except us. On best serp api, Claude drew 10 citations from our site out of only four source domains in the entire answer, and named SerpApi, DataForSEO, Serper.dev, Bright Data and Oxylabs instead.
There is one consistent exception, and it is not a happy one. Every time Google AI Mode did name us, the reason was price — "best low-cost alternative", "$10 risk-free entry point", "if price efficiency is your absolute highest priority". Four namings in this pass, no exceptions. The engines have filed us as the cheap option, not a capable one. That is a positioning problem, and it is visible in the data long before it is visible in revenue.
Repeatability: 16 of 46 answers moved
Before publishing any of this I ran the whole set again on a different day, because a single run of a non-deterministic system is a sample, not a measurement.
| Metric | 29 August | 31 August |
|---|---|---|
| Answers citing us | 30 | 26 |
| Answers naming us | 9 | 8 |
| Retrievals ending without a naming | 21 of 30 — 70% | 19 of 26 — 73% |
The aggregate is solid: 70% and 73%, and 71% when both runs are pooled into 56 retrievals.
The individual answers are not. Agreement between runs was 74% on citation and 89% on naming — 16 of the 46 answers flipped on at least one metric in two days. Queries that were completely absent in one run were cited first in the other.
So the honest reporting rule, which I am holding myself to here: publish the aggregate with its range, never a single cell. "Roughly 7 in 10 retrievals end without a naming" survives replication. "Engine X recommends us for query Y" does not, and anyone can falsify it in one try. That is the same conclusion the earlier citation repeatability test reached from a different angle, and it is why a single-run visibility score should never be treated as a fact.
What this means for your SEO team
1. Track two numbers, not one. Cited and named. If your tool only reports citations, it is reporting the easier half. Our own AI search visibility metrics breakdown covers what else is worth logging per answer.
2. Strip anchors before you brand-match. Otherwise citation chips inflate your score. This is the single most common measurement bug in this space, and it always errs in the flattering direction.
3. Kill personalisation before you measure. Temporary chat on ChatGPT, fresh sessions everywhere else. A personalised chat will happily tell you that you are the market leader.
4. Read the gap column as a positioning report. A large negative gap — heavily cited, rarely named — usually means your content is good enough to be evidence but your brand is not yet the answer. A large positive gap means the opposite, and it is worth a lot more.
5. Watch what you are named for. Being named only in price sentences is a real finding about how a category has filed you, and it is more actionable than the raw count.
Score your own answers (Python)
Two pieces. First, the split that everything else depends on — separate the citations from the prose before counting brands. Standard library only.
import re
ANCHOR = re.compile(r"<a\b[^>]*>.*?</a>", re.I | re.S)
HREF = re.compile(r'<a\b[^>]*href="(https?://[^"]+)"', re.I)
TAG = re.compile(r"<[^>]+>")
def split_answer(html: str):
"""Return (prose_without_citations, cited_hosts) for one AI answer."""
hosts = []
for url in HREF.findall(html):
host = url.split("/")[2]
host = host[4:] if host.startswith("www.") else host
if host not in hosts:
hosts.append(host)
prose = TAG.sub(" ", ANCHOR.sub(" ", html)) # anchors go FIRST, whole
return re.sub(r"\s+", " ", prose).strip(), hosts
def score(html: str, brand: re.Pattern, domain: str):
prose, hosts = split_answer(html)
return {"cited": domain in hosts,
"named": bool(brand.search(prose)),
"sources": len(hosts)}
Point it at a captured answer and the two metrics come out separately. The ordering matters: remove the anchor elements whole, then strip the remaining tags. Strip tags first and the chip text survives as prose, which is exactly the false positive this is meant to prevent.
Second, the index itself. This runs against the published dataset and reproduces the headline numbers in this article:
import json, urllib.request
URL = "https://apiserpent.com/data/ai-citation-index-2026.json"
data = json.load(urllib.request.urlopen(URL))
answers = data["answers"]
retrieved = sum(a["target_cited"] for a in answers)
recommended = sum(a["target_named"] for a in answers)
# Count the overlap directly. `retrieved - recommended` happens to give the same
# answer on this sample only because every answer that named us also cited us, and
# nothing guarantees that: an engine can name a brand it never linked to (eight
# brands in this index are named without a single citation). Subtracting would
# quietly under-report the gap the first time that happens to us too.
cited_not_named = sum(1 for a in answers if a["target_cited"] and not a["target_named"])
print(f"answers analysed {len(answers)}")
print(f"cited as a source {retrieved}")
print(f"named in the answer {recommended}")
print(f"cited but not named {cited_not_named}"
f" ({100 * cited_not_named / retrieved:.0f}% of retrievals)")
for row in data["leaderboard"][:6]:
print(f" {row['domain']:<18} cited {row['answers_citing']:>2}"
f" named {row['answers_naming']:>2} gap {row['gap']:+d}")
Output:
answers analysed 46
cited as a source 26
named in the answer 8
cited but not named 19 (73% of retrievals)
serpapi.com cited 22 named 44 gap +22
dataforseo.com cited 14 named 35 gap +21
serper.dev cited 7 named 35 gap +28
brightdata.com cited 9 named 29 gap +20
apiserpent.com cited 26 named 8 gap -18
scrapingdog.com cited 13 named 18 gap +5
To run this continuously against your own domain rather than by hand, the AI Mode citation tracker and the Python citation tracker both plug into Serpent API and give you the capture half; the scoring above is the part they leave out.
Limitations
Stated plainly, because the numbers are only worth what the method is worth.
- n = 12 queries. One vertical, one commercial intent. This is a field measurement, not a census.
- One day, one geography, one browser profile. Captured 31 August 2026. AI answers vary by location and account.
- Gemini is not in the citation index. It exposes no source list, so citation cannot be measured there at all. It contributes to naming counts only.
- Perplexity contributed 10 answers, not 12, in the second pass — its free tier's daily search limit cut the run short. Both missing questions were absent for us in the first pass too.
- "Named" is a brand match, not a sentiment read. Two of the matches in the wider capture were not recommendations at all: one warned readers off a group of low-cost providers including us, and one credited us as the publisher of a comparison. The dataset flags mention type; a naive count would have scored both as wins.
- I am a vendor in my own study. See the disclosure. Every input is published so you can check the arithmetic.
FAQ
What is the difference between being cited and being recommended by an AI engine?
A citation is a link the answer used as a source. A recommendation is the brand the answer actually names in its text. They are separate lists and they barely overlap. In this study of 46 AI answers, 121 domains were cited but only 28 brands were named, and 98 of the cited domains were never named once.
Does getting cited by AI search help my brand?
Less than most dashboards imply. Across 46 answers the domain that was cited most often was named in only 9 of them, while a brand that was never cited once was named in 18. A citation puts your page in the model's working set for that answer; it does not put your name in the sentence a reader acts on.
How many sources does an AI answer actually cite?
It varies hugely by engine. In this sample the median was 9 distinct domains for Google AI Mode, 10 for Perplexity, 9 for Claude and 2 for ChatGPT. ChatGPT names roughly as many brands as the others while linking to a fraction of the sources.
Are AI citation results repeatable?
Not at the level of a single query. The same 12 questions were run twice, two days apart. The aggregate barely moved, from 70% to 73% of retrievals ending without a naming, but 16 of the 46 individual answers changed on at least one metric. Treat a one-run AI visibility score as a sample, not a measurement.
How do I measure cited versus recommended myself?
Capture the rendered answer, then split it into two things before you count anything: every outbound link host, and the answer text with all anchors and citation chips removed. Brand names inside citation chips are sources, not recommendations, so counting them turns a citation into a false positive. The Python in this article does exactly that.
Which AI engine names the most brands per answer?
Claude, at an average of 7.2 named brands per answer in this sample, ahead of Perplexity at 5.9, Google AI Mode at 5.3 and ChatGPT at 5.1. Claude also cited the study's publisher on 11 of 12 questions while recommending it on none of them.
Measure it yourself
Serpent API returns Google, News, Images, Shopping and Maps results as clean JSON — enough to run a citation index like this one every week instead of once. Pay as you go, and deposited credits do not expire. Rates start at $0.03 per 1,000 calls on the Scale tier, which is locked in by a one-time $500 deposit; the entry tiers cost more per call and need no deposit.



