How Much Do Google, Bing and Brave Results Overlap?

By Serpent API Team · · 17 min read

Ask Google how much search results overlap and it will tell you, in its own AI Overview, that around 84.9% of results appear on only one engine. That figure is real. It is also from 2006, and it describes Ask Jeeves, MSN Search, Google and Yahoo — three of which no longer exist in that form.

That is the actual state of this question. The problem is not that nobody has measured how much search engine results overlap. It is that the most visible answer in 2026 is twenty years old, and almost nothing has been published since that is both independent and large enough to replace it.

This post is two things. First a survey: every study we could find, assembled into one table — sample size, date, engines, method, and who paid for it — because that table does not exist anywhere else, and because the differences between those studies explain almost all of the apparent disagreement in the field.

Then an experiment. The gap in this field is not analysis, it is a current dataset with its method stated, so on 13 August 2026 we ran one: the same 300 queries through Google and Brave, back to back, US market, top 10 each. Those results are in their own section below, and our row in the table is labelled vendor-published like every other 2026 row, because that is what it is.

TL;DR: The largest published study — 12,570 queries, Spink et al. (2006) — found 84.9% of results appeared on only one of four engines, and just 1.1% on all four. Google's own AI Overview still serves that 2006 figure as the current answer. The best modern academic number puts Google–Bing overlap under 32% (Yagci et al., 2022). We also ran our own collection: 298 usable Google–Brave query pairs on 13 August 2026, sharing 42% of top-10 URLs under a union denominator, 58% under a min denominator, and 54% / 74% when matching on registrable domains instead. One collection, four honest numbers, spanning 42% to 74% — which is the whole problem with every figure on this page, ours included. Every number that is not ours belongs to someone else and is linked.

Every published overlap study, in one table

Here is the whole field, as far as we could trace it. The column that matters most is the last one: whether the people who published the numbers also sell something that the numbers reflect on.

StudyDateQueriesEnginesHeadline findingPublished by
Spink, Jansen, Blakely & Koshman — Information Processing & Management 42(5)200612,570Ask Jeeves, Google, MSN, Yahoo84.9% of results unique to one engine; 1.1% shared by all four; only 7% of first results ranked alikeAcademic
Bar-Ilan, Mat-Hassan & Levene — Computer Networks 50(10)2006Methods paper: compares five ways of scoring rank similarity and proposes a sixthAcademic
Rather, Lone & Shah — Library Philosophy and Practice (e-journal), paper 226200820Google, AltaVista, HotBot, Scirus, BiowebHotBot overlapped most with the others, followed by Google; compound and complex queries overlapped more than simple onesAcademic
Lewandowski — JASIST 66(9)20151,000 + 1,000Google, BingRelevance, not overlap. Navigational queries answered correctly by Google 95.3% vs Bing 76.6%Academic
Yagci, Sünkler, Häußler & Lewandowski — ASIS&T 85th20223,537Google, Bing, DuckDuckGo, MetaGerGoogle–Bing overlap always under 32%; MetaGer–Bing up to 78%Academic
RankCaster overlap analysisApr 202645 platformsToo small to generalise from; listed for completenessVendor
Vizup — AI Search Results OverlapJul 2026AI answer enginesNo original data. A synthesis of third-party figures, carrying its own "verify before publishing" noteVendor
Crawlora — SERP Volatility IndexJul 202660Google, Bing, BraveGoogle–Bing 31.3%; Google–Brave 71.3%; 58.1% of domains on exactly one engine; #1 agreement 24.6%Vendor (CC BY 4.0)
Serpent API — Google–Brave overlap run (this post)Aug 2026298Google, BraveExact URL 42% union-normalised / 58% min-normalised; registrable domain 54% / 74%; RBO 0.563–0.647 on URLs. Collection modes mixed, declaredVendor — that is us. We sell a SERP API, so this row reflects on something we sell. Read it with the suspicion the column is for

Two things jump out. The independent academic work is old — the most recent substantial study predates AI Overviews entirely. And every 2026 row, including the one we added ourselves, comes from a vendor.

The 2008 entry deserves a note, because its title oversells it. "Overlap in Web Search Results: A Study of Five Search Engines" reads like a peer to Spink, and it gets cited as one. Open it and the sample is 20 queries, all in biotechnology, across five engines of which two — Scirus and Bioweb — are specialist science search tools rather than general web engines, and three no longer exist at all. It is a real paper, and its finding that compound queries overlap more than simple ones is genuinely interesting. It is not a general-web overlap study, and citing it as one would be a worse error than leaving it out.

We include it because it makes the point better than any argument could: the published record on this question is thinner than its citation counts suggest. Two of the entries above are vendor marketing, one is a synthesis with no data of its own, and one of the two academic papers that names five engines is twenty biotech queries.

That is the honest state of the field, and adding a row of our own does not change it: there is still no independent 2026 measurement to point at. Anyone quoting a current cross-engine overlap number to a decimal place is quoting one of the vendor rows at the bottom of that table — so read every one of them, ours first, with its method attached rather than as a settled figure.

Google is answering this question with a 2006 study

Here is the clearest illustration of the gap, and you can check it yourself in about ten seconds.

Observed 13 August 2026: searching Google for search engine results overlap study returned an AI Overview stating that "roughly 85% of results appear on only one major search engine" and that "around 84.9% of results on a first results page are unique to a single search engine" — sourced to ScienceDirect. That is Spink et al., published in 2006, measuring Ask Jeeves, MSN Search, Google and Yahoo.

Google is answering a 2026 question with a 2006 number about search engines that mostly no longer exist, and presenting it as settled. Not because the study was bad — Spink et al. is the largest and most careful work in the field — but because there is so little newer to cite. That is the hole our own collection was run into, and one vendor-published collection does not fill it.

Two caveats, in the spirit of the rest of this post. AI Overviews are generated per query and vary by person, location and time of day, so what you see may differ from what we saw; run the query and find out, because that variability is itself part of the story. And the underlying Spink figures are not in doubt — we checked them against the paper, and they are in the table above with the citation.

Only three of these engines have their own index

Before comparing any two engines, it's worth knowing whether they are actually two engines. Among the names that appear in overlap studies, only Google, Bing and Brave crawl and rank from an index they built themselves.

This reframes the Yagci finding. Their headline is that Google and Bing overlap under 32% — but the same study found MetaGer and Bing reaching 78%. That second number is not two engines agreeing. It is one index arriving through two front doors.

So the useful distinction in any overlap table is within-family versus cross-family. Within-family figures (Bing with DuckDuckGo, or Bing with the Yahoo/Bing front end) quantify how far a syndication partner drifts from its source through re-ranking and filtering. Cross-family figures (Google with Bing, Google with Brave) are the only ones measuring genuine algorithmic disagreement. Reporting them in one column, as most roundups do, mixes two different quantities.

The headline numbers, and why they disagree

Line up the two most-cited Google–Bing figures and they look reassuringly consistent: under 32% from Yagci et al. in 2022, and 31.3% from Crawlora in July 2026. Fourteen years apart, and the two figures land within a point of each other.

They are not comparable, and the reason is the denominator. Crawlora states plainly that its 31.3% is "of the shorter result set." Divide shared domains by the shorter of two lists and you get a larger number than dividing by their union, which is the convention in the academic work. The two studies agreeing to within a point is a coincidence of two different fractions, not a replication.

The figure that should raise an eyebrow is Crawlora's 71.3% Google–Brave overlap — more than double its own Google–Bing number. Brave runs a genuinely independent index, so on the face of it Brave should look less like Google than Bing does, not more.

The first thing to check is whether it is even the same kind of number as the academic figures it gets quoted beside. It is not. Crawlora states its method plainly: matching is on registrable domains, and overlap is reported as a share of the shorter result set. Two engines that return nytimes.com on different articles count as agreeing. Set that against Yagci's under-32%, computed on results rather than bare domains, and the two numbers are not measuring the same thing. Placing them side by side, as several roundups now do, is an apples-to-oranges comparison.

Within Crawlora's own consistent method the gap survives — Brave 71.3% against Bing 31.3%, both scored the same way — and that is the genuinely interesting result. It is also resting on 60 keywords, where a handful of queries can move a mean several points.

We flag it because it is the most quotable number in the 2026 literature and the one most likely to be repeated without its caveats. If you are citing it, cite the matching unit and the denominator with it, or you are not citing the same finding. Our write-up of Brave versus Google as SERP data sources goes into where the two genuinely diverge. The next section is what happened when we ran a Google–Brave collection ourselves and reported it every legitimate way at once.

Our own run: 298 Google–Brave query pairs

Everything above this line belongs to someone else. This section does not.

On 13 August 2026 we put the same 300 queries through Google and Brave back to back, both requested for the US market, and truncated both result sets to the top 10 before any comparison. 298 pairs were usable. Then we computed overlap four ways — two denominators × two matching units — on one set of captures.

The short answer: Google and Brave shared 42% of their top-10 URLs when the shared count is divided by the union of both lists, and 58% when it is divided by the shorter list. Match on registrable domain instead of exact URL and the same captures read 54% and 74%. Same lists, same moment, four legitimate numbers.

Matching unitunion-normalised (Jaccard)min-normalised (overlap coefficient)RBO_min / RBO_ext, d=10, φ=0.9
Exact canonical URL42% (95% CI 40–44)58% (95% CI 56–59)0.563 / 0.647
Registrable domain54% (95% CI 52–56)74% (95% CI 72–76)0.599 / 0.742
One Google-Brave collection, reported four ways Horizontal bars showing overlap at depth 10 from the same 298 query pairs: exact URL union-normalised 42 percent, exact URL min-normalised 58 percent, registrable domain union-normalised 54 percent, registrable domain min-normalised 74 percent. The same 298 query pairs, reported four ways Exact URL · union-norm. Exact URL · min-norm. Domain · union-norm. Domain · min-norm. 42% 58% 54% 74% 0% 100%
Stratum A, n=278, collected 13 August 2026. Nothing changes between the bars except the matching unit and the denominator.

Stratum A, n=278. Stratum B (n=20) separately: exact URL 40% (CI 31–50) union-normalised and 55% (CI 45–64) min-normalised; registrable domain 51% (CI 43–59) and 71% (CI 64–77). At n=20 no Stratum B mean stands on its own — it is shown because the pre-registration says both strata get reported.

How the 300 queries were chosen

The sample is in two declared strata. Stratum A (278 usable) is a seeded random draw from a public Wikimedia frame: the 1,000 most-viewed English Wikipedia articles for a fixed date, minus administrative pages and titles under three characters. One published integer regenerates the entire list and the collection order, so the selection is not something you have to take our word for.

Stratum B (20) is hand-built and intent-stratified — informational, commercial, navigational, technical — and fixed in writing before anything was collected. Stratum A exists precisely because a hand-picked sample chosen by people who already know the hypothesis is not a pre-registered sample. Comparing the two is itself one of the findings.

Two of the 300 failed to collect and are excluded pairwise: one upstream HTTP 503 on the Google side, and one row that was not actually served as US, which the market control caught and excluded. A row served for a different market measures country and would report it as engine. No result list came back short of 10.

Finding 1: the matching unit moves the headline by 12 to 16 points

This is the number worth taking away. Going from exact canonical URL to registrable domain, on identical captures, moves union-normalised overlap from 42% to 54% (12 points) and min-normalised overlap from 58% to 74% (16 points). Add the choice of denominator and one collection can be reported anywhere from 42% to 74% with nobody misreporting anything.

That range is wider than the distance between most of the published figures people argue about. It means a cross-engine overlap percentage quoted without its matching unit and its denominator is not a weak number — it is an uninterpretable one, and no honest comparison to another study can be built on it.

It also puts Crawlora's 71.3% Google–Brave figure in context. That is a domain-level, min-normalised kind of number, and read the same way our collection gives 74% (95% CI 72–76) — within a few points. Read our captures the strict way instead, exact URL and union-normalised, and the same data says 42%. That is a statement about how their headline should be read, not a measurement of their study: we did not collect their keywords, their week or their capture paths, and nothing here confirms or refutes their result.

Finding 2: we pre-registered a prediction and it failed

We wrote down, before collecting, that Stratum A would run high — Wikipedia article titles as queries should inflate Wikipedia's presence in both engines' results and drag the two lists together.

It did not happen. On exact-URL union-normalised overlap, Stratum A came out 2 points above Stratum B (42% vs 40%), and Stratum B's interval (31–50%) comfortably contains Stratum A's mean. At n=20 Stratum B settles nothing by itself, so the honest reading is narrow: this run found no evidence that where the queries came from changes cross-engine overlap.

We are reporting it because it was pre-registered. A prediction you only publish when it works is not a prediction, and the query-provenance variable is one the vendor panels leave undeclared entirely — this suggests it may matter less than we assumed.

Where this run breaks our own rule

The third trap below is mixed collection modes, and it is the sharpest criticism this post makes of the 2026 panels. Our own collection breaks it. Google was collected through a third-party SERP API; Brave was collected by our own scrape. Two capture paths, on exactly the rule we tell everyone else to hold constant.

We are not going to soften that rule to accommodate our data. The rule is right, we did not meet it, and every number in this section is weaker for it — some unknown share of the disagreement we measured is the two paths rather than the two indexes. Rewriting our own standard because our results tripped over it is precisely the behaviour this post exists to criticise, so the standard stays where it is and the failure is ours.

What we can do is make the flaw priceable instead of invisible. The capture path is stated here, in the dataset's method file, and in every row of the raw captures. The query sample regenerates from one published seed. The script that turns the captures into the table above is released with them. That is the standard this post asks for — publish the sample and publish the raw data — and it is the one place we can point to a concrete difference: the 2026 vendor studies above do not name their queries at all, so nobody outside those companies can regenerate their sample, check it for selection effects, or recompute a single number in it.

What we are releasing

The bundle is the 300-query sample with its frame URL, seed and exclusion counts; every raw SERP exactly as each engine returned it; the sampler that regenerates the query list; the metrics script that produces every number in this section; and the dated public-suffix-list snapshot used to resolve registrable domains. URLs are stored raw on purpose — canonicalisation happens at analysis time, so you can change the matching rule and re-run the whole study without collecting anything again.

Both scripts ship with hand-computed self-tests and neither needs an API key. Email info@apiserpent.com for the bundle. If you find an arithmetic error in it, we would rather hear about it than not.

Four traps that make overlap numbers wrong

These are the reasons two competent people measuring the same thing get different answers.

1. The denominator

"Overlap" is not one metric. Take the same ten-result lists and you can divide the shared items by the union of both lists — that is Jaccard, the conservative reading — or by the shorter of the two, which is the overlap coefficient and is systematically higher. Same numerator, different denominator, two very different headline percentages.

Worth being precise here, because it trips people up: union-normalised overlap@10 is Jaccard. They are not two metrics that can corroborate each other, and reporting both as though they were is double-counting one number. Report the pair — union-normalised and min-normalised — and name each.

2. The matching unit

Are two results "the same" when the URLs match, or when the domains match? Domain-level matching counts two different nytimes.com articles as agreement, and it produces dramatically higher numbers than URL-level matching on identical data. Our own captures put that gap at 12 to 16 points, depending on the denominator. Crawlora's 71.3% is a domain-level figure; most of the academic work is stricter. A study that does not state its matching unit has published an uninterpretable number.

There is a mechanical trap underneath this one. Engines wrap their outbound links: Bing serves bing.com/ck/a?…&u=a1<base64> redirect wrappers and Google serves /url?q=…&ved=…. Match the raw href attributes and cross-engine URL overlap collapses toward zero as a pure link-extraction artefact, with nothing to do with the engines. Any URL-level comparison needs a canonicalisation step first: decode the wrappers, strip tracking parameters, normalise scheme, host and trailing slash, and resolve the registrable domain against a real public-suffix list rather than by splitting on dots — otherwise bbc.co.uk and nhs.uk come out wrong.

3. Mixed collection modes

This is the sharpest problem in the 2026 data. Crawlora's panel states that Google was captured from browser sessions while Bing and Brave came from API endpoints. API responses and rendered SERPs are different artefacts — different result counts, different inclusion of ads and universal blocks, sometimes different ranking. A comparison built across two capture paths is partly measuring the paths. Hold the collection mode constant across every engine, and publish the raw captures so a reader can check the arithmetic themselves rather than take the summary on trust.

That rule indicts our own collection as well — we captured Google and Brave down two different paths. The rule is printed here exactly as it was written, not trimmed to fit our data.

4. Timing and drift

Search results are not stable enough to sample casually. Results move with location, time of day, personalisation state, and live experiments — we walk through the six mechanisms in why Google results differ every single time. If engine A is sampled in the morning and engine B in the afternoon, the difference between them includes a day's drift. Every engine, every keyword, inside a window of minutes — otherwise drift is confounded with disagreement. The 2026 vendor panels do not state their collection timing at all; ours carries both timestamps in every row, and the median gap between the two engines on a query was about six seconds.

How to measure it yourself

The gap in this field is not analysis, it is a current, method-clean dataset. Here is the design we used above, so you can audit it or run your own. Nothing in it is proprietary — it is assembled from the methods literature cited at the foot of this post — and where our own run failed to meet a point, that is noted against the point rather than quietly dropped.

If you are wiring up the collection side of that, our walkthrough of building a multi-engine search aggregator covers parallel querying, result deduplication and cross-engine consensus scoring in Node.js.

What this run does not settle

We ran one collection. It is worth being exact about what it can and cannot carry.

Two engines, one country, one day. Google and Brave, US-served results, 13 August 2026. It says nothing about Bing, nothing about DuckDuckGo, nothing about non-US markets, and nothing about whether these figures hold next month — results move constantly, for the six reasons we walk through separately. Overlap measured in one window is a snapshot, not a constant.

Two capture paths, not one. The mixed collection mode is the single biggest weakness in the numbers above, and it is our own rule that we broke. It is disclosed everywhere the data appears; disclosure is not a fix, it is just the difference between a limitation you can weigh and one you cannot see.

The queries are not a general-web distribution. 278 of the 298 pairs come from Wikipedia article titles. People really do search those things, but that is not what all searching looks like. Finding 2 suggests it mattered less than we predicted; it does not prove it did not matter.

And it is vendor-published. We run a SERP API, so a number showing engines disagreeing is a number that flatters us. That is exactly the objection this post levels at the other 2026 rows, and putting our own row in the table does not exempt us from it. The only answer we have is the one we asked of them: the sample, the raw captures and the code that turns one into the other are released, so nothing here has to be taken on trust.

What the run does not change is the shape of the answer, which has pointed one direction for twenty years. A single engine is a sample, not a census. On the most generous reading of our own data — domains, min-normalised, 74% — roughly a quarter of the shorter engine's list is absent from the other. On the strictest reading — exact URLs, union denominator — the two engines agreed on 42% of the distinct URLs they returned between them, so the remaining 58% appeared on one engine and not the other: the same shape Spink et al. found in 2006, on different engines, twenty years apart. (That 58% is a different quantity from the 58% min-normalised figure in the table above. The two coincide by arithmetic accident, not because they measure the same thing — which is itself a small demonstration of why a bare percentage on this topic is uninterpretable without its denominator.) If your product depends on knowing what "the web says" about a query, one source will systematically under-report it.

Work with more than one source of truth

Serpent's SERP API returns organic results, People Also Ask, AI Overviews and more as structured JSON, so you can pull comparable data across engines without maintaining a scraper for each one. 10 free searches, then from $0.60 per 1,000.

Get Your Free API Key

Explore: Google SERP API · Brave Search API · Playground · Docs

FAQ

Do different search engines return the same results?

No, and the gap is much wider than most people expect. The largest published study of the question, Spink et al. (2006), took 12,570 queries across Ask Jeeves, Google, MSN Search and Yahoo and found 84.9% of results appeared on only one of the four engines. Just 1.1% were shared by all four. Twenty years of newer work has not overturned that shape: Yagci et al. (2022) measured Google–Bing overlap at always under 32% across 3,537 queries, and the 2026 vendor panels report similar or lower figures. Search engines agree on far less than their reputations suggest.

Is the 84.9% figure still accurate in 2026?

Honestly, nobody knows. It is the most cited number in the field and it comes from Spink et al. (2006), which measured Ask Jeeves, MSN Search, Google and Yahoo across 12,570 queries. Three of those four engines no longer exist in that form, the web is dominated by different sites now, and AI answers sit above the organic list in a way they did not in 2006. The figure has never been refuted, but it has also never been replicated at that scale. The closest modern work, Yagci et al. (2022), found Google and Bing overlapping at always under 32% across 3,537 queries, which is consistent with the 2006 picture without confirming the specific number. Treat 84.9% as the best available historical estimate rather than a current measurement.

What is the overlap between Google and Bing search results?

The best independent measurement is Yagci et al. (2022), which found the overlap between Google and Bing was always under 32% on the top 10 results across 3,537 queries in Germany and the US. The most recent figure comes from Crawlora's July 2026 panel, which reported 31.3% page-one overlap on 60 keywords. Those two numbers agree closely, but they are not directly comparable: Crawlora normalises by the shorter of the two result sets rather than by their union, which inflates the percentage relative to the academic convention.

How much do Google and Brave results overlap?

We measured it on 13 August 2026. Across 298 query pairs, Google and Brave shared 42% of their top 10 URLs when the shared count is divided by the union of both lists, and 58% when it is divided by the shorter list. Match on registrable domain instead of exact URL and the same collection reads 54% and 74%. Rank-biased overlap on URLs is 0.563 to 0.647 at depth 10 with φ = 0.9. One collection, four legitimate numbers, spanning 42% to 74% — which is why an overlap figure quoted without its matching unit and its denominator cannot be compared with any other figure. One caveat we state everywhere: Google was collected through a third-party SERP API and Brave by our own scrape, so the collection mode was not held constant, and the result is weaker for it.

How many search engines have their own index?

Very few. Among the engines usually named in Western overlap studies, only Google, Bing and Brave crawl and rank from their own index. Yahoo/Bing is one index under two brands, because Yahoo's web results have been served from Bing for years, and DuckDuckGo states in its own documentation that it largely sources traditional links and images from Bing, supplemented by its own DuckDuckBot crawler and other specialist sources. This matters for any overlap study: high agreement across Bing, Yahoo/Bing and DuckDuckGo is mostly syndication of one index, not three engines independently reaching the same conclusion.

Why do published SERP overlap studies disagree with each other?

Four reasons, and all of them are method rather than substance. First, the denominator: some studies divide shared results by the union of both result sets, others by the shorter set, which can move the same raw data by ten points or more. Second, the matching unit: counting agreement at the domain level rather than the URL level produces much higher numbers on identical data, and Crawlora's 71.3% is a domain-level figure while most academic work is stricter. Third, collection mode: Crawlora's 2026 panel captured Google from browser sessions but Bing and Brave from API endpoints, and API and browser SERPs are not the same artefact. Fourth, timing: search results drift by the hour and by location, so any study that does not sample all engines inside a tight window is measuring drift alongside disagreement.

What is rank-biased overlap and why use it instead of a percentage?

Rank-biased overlap (RBO) is a similarity measure published by Webber, Moffat and Zobel in ACM Transactions on Information Systems in 2010. A plain overlap percentage treats a result at position 1 and a result at position 10 as equally important, and it cannot handle lists that share only some of their items. RBO is top-weighted, works on rankings that are only partly conjoint, and can be computed from a prefix of each list, which is exactly the cross-engine case. Report a plain overlap percentage for comparability with older studies, then report RBO as the rigorous figure.