# Gemini vs ChatGPT: 42% Overlap in the Brands They Recommend

I compared 38,966 questions put to both Gemini and ChatGPT. The brands they consistently recommend overlap just 42% of the time. Here is what that means.

**Published:** September 22, 2026
**Author:** Connor Lahey

---

</KeyTakeaways>

Most people comparing Gemini and ChatGPT want to know which one to open tomorrow morning. On everyday work the two have converged far enough that most people won't reliably notice the difference, and the right pick mostly comes down to which tools you already pay for.

For marketers, there's another question worth asking: what happens when people use these tools to find products and services? Ask either engine for the best project management tool, or the best accounting software in Manchester, and it names a handful of brands. Whether it names yours is one question. Whether it names yours consistently across Gemini and ChatGPT is another. Until now, there's been little data on how much those answers actually differ.

I analyzed 38,966 questions asked repeatedly of both engines between July 8 and September 15, 2026 in Searchable internal data, then compared the brands each one returned. The two engines agreed far less often with each other than either agreed with itself. If you track one engine as a proxy for the other, you're missing a substantial part of the picture.

## How Gemini and ChatGPT differ as answer engines

Ask both engines the same question and you get two answers that look alike. Similar length, similar confidence, a similar handful of brands. The differences show up underneath, in how each one builds the answer and what it reads to get there.

Those differences are what this study measures. Here is the shape of them before the numbers that follow.

| | Gemini | ChatGPT |
|---|---|---|
| Cites at least one source | 69% of answers | 96% of answers |
| Sources listed per answer | Fewer | More than twice as many |
| Reads most from | Competitor pages and a brand's own site | Institutional pages and review sites |
| Brands named per question | 3.69 | 3.73 |
| Agreement with its own earlier answers | 75.0% | 76.8% |

*Searchable internal data. Citation rows cover August 23 to September 15, 2026; the rest cover July 8 to September 15, 2026. See the method for why the citation figures use a shorter window.*

### What the two engines do the same way

Both search the live web before answering, and both settle on a similar number of brands. Gemini names 3.69 per question on average, ChatGPT 3.73. Neither is scattering names around while the other stays selective.

Both are also about as consistent with themselves. Split either engine's answers to the same question into two batches and the batches agree around three quarters of the time. That matters more than it sounds, because it sets the bar for everything that follows: it's how much agreement you should expect from an engine that hasn't changed its mind.

### Where they pull apart

ChatGPT shows its working. It names at least one source in 96% of answers against Gemini's 69%, and lists more than twice as many sources when it does. For anyone hoping to be one of those sources, that's a wider door.

The two also read different corners of the web. Gemini leans on competitor pages and on a brand's own site. ChatGPT leans on institutional pages and review sites. Same question, different reading list.

That citation difference is an observed behavior in this dataset, not evidence that ChatGPT is inherently better at research or that its answers are more accurate.

## How often do Gemini and ChatGPT name the same brands?

Across 38,966 question sets put to both engines between July 8 and September 15, 2026, the brands Gemini consistently recommends and the brands ChatGPT consistently recommends overlapped just 42% of the time. Being recommended by one engine tells you very little about whether the other will name you at all.

This doesn't mean either engine is wrong 58% of the time. It means the sets of brands they consistently named overlapped by 42% under the study's definition of a recommendation.

<div className="not-prose my-8">
  
</div>

Broken down by direction, the picture is symmetrical. Of the brands ChatGPT reliably recommends, Gemini names 57.7%. Of the brands Gemini reliably recommends, ChatGPT names 57.6%. So more than four in ten of the brands each engine puts in front of a user are missing from the other's list.

That symmetry matters. If one engine were simply noisier, the disagreement would run mostly one way: the noisier engine would produce a longer, more scattered list, while the more selective one would rarely add brands of its own. That isn't what happens. The two engines produce similarly sized recommendation sets, but those sets only partly overlap.

Here's what that means in practice. You ask ChatGPT, "what's the best CRM for a small agency," and find your brand in the answer. It's tempting to call that AI visibility, but you've just measured your visibility in one engine and assumed it reflects your visibility in the other. That assumption is wrong about four times in ten, in both directions, including for the competitor you think you're beating.

A recommendation from one engine doesn't automatically carry over to another. It's more useful to treat visibility as platform-specific, much as you would in separate search environments. The work that earns you a place in one may also improve your chances in the other, but it doesn't reliably produce the same recommendations.

## How I compared Gemini and ChatGPT across 38,966 question sets

I identified 38,966 questions that were put to both Gemini and ChatGPT repeatedly over ten weeks, then compared the brands each engine named consistently. A brand counted only if an engine mentioned it in at least half of its answers to that question, which keeps the two lists comparable in length.

The data is Searchable internal data. The platform asks each tracked question on a schedule, engine by engine, and keeps every answer. I took every unbranded question that ran on both Gemini and ChatGPT, in the same locations, between July 8 and September 15, 2026: 38,966 questions, covering 1,410,364 ChatGPT answers and 1,040,107 Gemini answers. Every figure in this article is a snapshot of that dataset taken on September 18, 2026.

Then I ran the check that tells us whether the 42% means anything. AI answers vary between runs. If each engine only agreed with itself 45% of the time, 42% between two engines would be unremarkable. So I measured that first. I split each engine's own answers into two batches and compared the batches, using the identical method. Then I ran the opposite test: pooled every answer from both engines, threw away the engine label, and split at random.

### Checking the 42% against noise

<div className="not-prose my-8">
  
</div>

*Searchable internal data, July 8 to September 15, 2026. All four rows compare half-batches of answers, which is what makes the last row comparable to the first two. 22,424 question sets for the split-batch rows, 25,602 for the shuffle.*

<div className="not-prose my-8">
  
</div>

Each engine agrees with itself about three quarters of the time. Two different engines agree about four times in ten. Strip the engine label out and agreement climbs back to three quarters, which is what you'd expect if the label meant nothing.

The result is a substantial gap between within-engine consistency and cross-engine overlap. That's the reason a single-platform visibility check can't stand in for visibility across both platforms.

Three of my most dramatic numbers didn't survive testing and I threw them out. I explain those exclusions, the full method, and every limitation at the end of this article.

## Does ChatGPT or Gemini cite more sources?

ChatGPT cites at least one source in 96% of its answers, against 69% for Gemini. It also lists far more sources per answer, more than twice as many as Gemini in every week of the study. Figures cover August 23 to September 15, 2026, across every answer to the unbranded questions put to both engines.

<div className="not-prose my-8">
  
</div>

The denominator matters here, because numbers like this get quoted loosely. That 96% is the share of ChatGPT's answers carrying at least one source. It isn't the share of brands that get cited, and it isn't the odds of any one page being picked. Different studies pick different denominators and land on wildly different figures, so it's the first thing to check when somebody quotes you one, including this one.

I won't give you an exact number of sources per answer. That count doubled halfway through the study, and it was us who changed, not ChatGPT. We started recording more of the source list than we had been. What didn't change is the gap. ChatGPT listed more than twice as many sources as Gemini for the same question, every week, before and after. The count itself never sat still like that.

For a marketer, that gap is the part that matters. An engine that lists more sources for an answer gives your brand more chances to appear. Gemini lists far fewer sources for the same question, so there are fewer opportunities for your brand to be included. That helps explain why the same content strategy can produce different results across the two engines.

## Which sources do ChatGPT and Gemini each rely on?

ChatGPT and Gemini draw on different kinds of pages. Institutional sources are the clearest split: 19.4% of what ChatGPT cites against 3.7% for Gemini in these weeks, and never less than two and a half times Gemini's share at any point in the study. Review sites run a little under twice as often. Gemini goes the other way. Competitor pages make up 37.3% of its citations against ChatGPT's 26.5%, and it cites a brand's own website about 1.4 times as often.

| Source type | ChatGPT | Gemini |
|---|---|---|
| Competitor sites | 26.5% | 37.3% |
| All other types | 25.4% | 26.2% |
| Editorial | 22.6% | 26.9% |
| Institutional (government, academic and similar bodies) | 19.4% | 3.7% |
| Review sites | 3.3% | 1.9% |
| Brand's own site | 2.7% | 3.9% |

*Share of cited pages by type, summing to 100%. "All other types" covers third-party, forum, social and unclassified pages. Searchable internal data, August 23 to September 15, 2026, across every answer to the unbranded questions put to both engines.*

This is the table underneath the 42%. The engines aren't disagreeing at random. They're reading different rooms. The widest gap is institutional pages, which make up about a fifth of what ChatGPT cites and under a twentieth of what Gemini cites.

Editorial is the one row I wouldn't read anything into. The two engines sit close together here, and in the weeks before this window ChatGPT was slightly ahead instead. Every share in this table moves from week to week. The wide gaps are the finding. The exact numbers are a snapshot.

One limit on all of this. These are shares of what each engine cites, not evidence that any source type earns a recommendation. Nothing here shows that a review-site mention makes ChatGPT more likely to name you. It shows that review sites make up more of what ChatGPT cites. That's a weaker claim, and it's the one the data supports.

### What each engine's cited pages look like

<div className="not-prose my-8">
  
</div>

There's a difference between what an engine was trained on and what it goes and fetches the moment you ask. [Where ChatGPT gets its data](https://www.searchable.com/blog/where-does-chatgpt-get-its-data) covers the training side, and [where Google gets its information](https://www.searchable.com/blog/where-does-google-get-its-information) covers the knowledge-source side. This study was about the fetching side.

## How to get your brand recommended by Gemini and ChatGPT

The pages ChatGPT cites skew toward sources you don't own: independent reviews and authoritative institutional pages. Gemini skews toward pages you do own, plus the comparison content around your category. Running one plan for both and expecting the same outcome doesn't follow from what either engine reads.

That follows from the source mix above, not from any theory about how either engine works inside. I measured what they cite. I can't see into either system, and neither can anyone outside those two companies.

Treat what follows as where to look first, not as a guarantee. It's inferred from what each engine cites, not from a test that changed one thing and measured the result. It's also where [answer engine optimization](https://www.searchable.com/blog/what-is-aeo) departs from the SEO playbook most teams already run.

### Where to focus for ChatGPT

Third-party coverage is the bigger opportunity. Institutional pages make up several times the share of ChatGPT's cited pages that they do of Gemini's, and review platforms a little under twice, so the work sits closer to digital PR than to on-site optimization. Getting into the places that already get cited matters more than the shape of your own pages, and it's the practical route to [show up in ChatGPT](https://www.searchable.com/blog/how-do-i-show-up-in-chatgpt-simple) at all.

One thing to know before you start: [AI citations vs backlinks](https://www.searchable.com/blog/ai-citations-vs-backlinks) work differently. A page can cite you without linking to you. The mention is what carries you into the answer.

### Where to focus for Gemini

Your own site does more of the work here. Gemini cites brand-owned pages at about 1.4 times ChatGPT's rate, so clear, current pages that spell out what your company does get read more often.

Comparison content is the second opportunity. Competitor pages make up 37.3% of Gemini's cited mix against ChatGPT's 26.5%, which means that the roundups and "best X for Y" posts in your category are doing as much for your Gemini visibility as anything you publish yourself. Being named accurately on those pages is worth chasing.

## How to track whether Gemini and ChatGPT recommend your brand

You can't tell whether Gemini and ChatGPT recommend your brand by asking each one once, because their answers vary between runs. Reliable measurement means putting the same questions to each engine repeatedly and tracking how often your brand appears across those answers.

That variance is the whole problem with a manual check, and it's why I used about 30 answers per question per engine rather than just one. Searchable's guide to [AI visibility tracking](https://www.searchable.com/blog/ai-visibility-tracking) covers the metrics that hold up and the ones that don't.

Searchable is built for this. It runs your tracked questions against Gemini and ChatGPT, two of the nine AI platforms it covers, then reports a Visibility Score, your Share of Voice against the competitors you track, and your average position when you do get named. Results are broken out per platform, so you can see the split this study describes inside your own data rather than inferring it from mine.

## Frequently asked questions

## Methodology

I compared 38,966 unbranded questions that ran on both Gemini and ChatGPT between July 8 and September 15, 2026, across Searchable internal data. Dates are UTC. A brand here means one that a project tracks: its own, plus the competitors on its list, including ones Searchable identified automatically. It counts as recommended when an engine named it in at least half of its answers to that question. Every figure is a snapshot taken on September 18, 2026, and all of them will move.

Prospect audits and archived brands are excluded. The citation and source-type figures use a shorter window, August 23 to September 15, for the reason below. This is a tested finding from observational data, and an exploratory one.

### Why a brand has to appear in half the answers

ChatGPT answered each question about 36 times. Gemini answered about 27. Count every passing mention and ChatGPT gets a longer list just for taking more turns, and longer lists overlap less. The threshold evens that out: 3.73 brands per question against Gemini's 3.69.

Move the bar and the answer barely changes.

- Count a brand when an engine names it in at least **30%** of its answers, and the two engines agree on **41%** of the brands.
- Raise the bar to **50%**, which is what this study uses, and they agree on **42%**.
- At **70%**, agreement rises to **47%**.
- At **90%**, it reaches **54%**, but by then most lists have shrunk to a single brand and fewer than half the questions still qualify, which is why the figure climbs.

Every one of those sits well below the three quarters each engine manages against itself.

### Three more checks

**The brand matching works.** Run the same method on questions that already name a brand and the engines agree 84.6% of the time, median 100%.

**No single brand drives it.** Measured brand by brand instead of pooled the median is 41.4%, close to the pooled figure, and no one brand makes up a large enough share of the sample to move it.

**Thin questions aren't the cause.** Keep only questions with at least five answers from each engine and the overlap is 42.9%.

### The number moves with the window

Ten weeks gives each engine's list enough runs to settle. Shorter stretches inside the same period, where there are fewer answers per question, run between 37% and 41%. A different ten-week window, June 1 to August 9, returns 42% again.

So read 42% as a measurement over a stated window, not a fixed trait of the two engines. What stays put is the distance between cross-engine overlap and each engine's agreement with itself.

### Three results I threw away

Every one of them made the gap look bigger than it is, which is why none of them are above.

**A monthly decline.** On an earlier window the overlap looked like it was sliding month by month. Held to a fixed set of the same questions it barely moved, so most of the apparent drop came from the sample changing rather than the engines.

**Questions with nothing in common.** A share of question sets shared no brands at all, but most of those were questions where one engine named only one or two brands, so a single miss wiped out the overlap.

**An early, larger gap**, calculated before correcting for ChatGPT's extra runs.

### What this study can't tell you

- **42% is an upper bound.** I can only count brands that are already tracked, and both engines name others beyond that set. Gemini names more of them than ChatGPT does, so true agreement is probably lower than 42%, not higher.
- **This is tracked data, not the market.** These are the questions tracked in Searchable internal data, so they lean toward category and comparison shapes like "best CRM for a small team."
- **Engine versions are pooled.** I compared the two products as a user meets them, not two model releases.
- **Gemini's volume nearly doubled** at its mid-August peak. Each question is still compared against itself, so a growing sample changes how many questions there are, not how the engines answered them.
- **Our own source capture changed** partway through the study, on our side rather than ChatGPT's. That is why the citation and source figures start on August 23 and why I quote no per-answer source count.
- **Rewording is untested.** 38,966 distinct questions covers a lot of phrasings, but I didn't test the same question asked another way.
- **Source shares are composition, not cause.** They show what each engine cites, not what earns a recommendation.

---

[Back to Blog](https://www.searchable.com/blog) | [Searchable Homepage](https://www.searchable.com)
