We tested the same 24 B2B generative engine optimization scenarios across five AI platforms. That produced 120 prompt-platform evaluations, 266 agency recommendations and 379 citation rows.
The clearest result wasn’t a universal winner. It was fragmentation. No agency appeared in all five systems, and the overlap between any two platforms’ recommendation pools ranged from only 2.7% to 13.5%.
Short answer: Claneo had the strongest cross-platform visibility, appearing 24 times across four systems. Seowerk appeared 18 times across three systems, and Peak Ace appeared 17 times across three. But no agency was consistently recommended everywhere. A company can look highly visible in one AI product and remain absent from another.
What we tested
We created 24 English-language prompts around B2B GEO in Germany and Europe. Twelve asked the systems to recommend agencies for specific commercial situations, including SaaS, fintech, crypto, cybersecurity, international expansion and Bavaria-based companies. The other twelve covered measurement, schema, technical SEO, AI referrals, prompt monitoring, content formats and a 90-day GEO plan.
Each scenario was processed by ChatGPT, Gemini, Perplexity, Claude and Microsoft Copilot on September 22, 2026. The systems were instructed to treat every scenario as an independent request, preserve recommendation order, report whether web search was used and return the sources used for that answer.
The headline uses “120 prompts” for readability, but these were not 120 different questions. The exact design was 24 unique scenarios × 5 platforms = 120 evaluations.
Platforms and modes recorded during the test
Platform
Model label shown
Mode
Reported web-search use
ChatGPT
GPT-5.6 Sol
Not shown
24 of 24
Gemini
Gemini
Default
14 of 24
Perplexity
Perplexity
Pro
24 of 24
Claude
Claude Opus 5
Not shown
21 of 24
Microsoft Copilot
Copilot
Not shown
24 of 24
The location shown or inferred by the platforms was Bayreuth, Bavaria. This describes the test environment, not the location of the Optis Digital office. Several prompts separately specified Munich, Bavaria, Germany or the wider DACH market.
The main finding: AI recommendation sets barely overlapped
If AI recommendations worked like a stable directory, the same established agencies would keep appearing. That isn’t what we found.
The largest pairwise overlap was between Perplexity and Claude. They shared five recommended entities, equal to a 13.5% Jaccard overlap between their combined recommendation pools. ChatGPT and Gemini shared only one entity, producing a 2.7% overlap.
Platform pair
Shared agencies
Pool overlap
ChatGPT vs Gemini
1
2.7%
ChatGPT vs Perplexity
2
3.9%
ChatGPT vs Claude
3
7.1%
ChatGPT vs Copilot
3
7.1%
Gemini vs Perplexity
3
9.4%
Gemini vs Claude
1
3.8%
Gemini vs Copilot
3
12.5%
Perplexity vs Claude
5
13.5%
Perplexity vs Copilot
3
7.7%
Claude vs Copilot
1
3.0%
This matters for measurement. A “top three” position in one chatbot doesn’t establish broad AI visibility. A useful benchmark has to monitor several systems with the same prompt set and track the result over time.
Which agencies appeared most often?
We counted an agency once each time it appeared in a prompt’s ordered recommendation list. The 60 commercial evaluations produced 266 recommendation slots. “Coverage” below means the percentage of those 60 commercial prompt-platform evaluations in which the agency appeared.
Agency
Recommendations
Platforms
#1 positions
Average position
Commercial coverage
Claneo
24
4
5
2.67
40.0%
Seowerk
18
3
6
2.17
30.0%
Peak Ace
17
3
9
2.29
28.3%
Seokratie
10
2
5
2.10
16.7%
pageworkers
10
1
3
2.70
16.7%
Suxeedo
9
2
1
3.11
15.0%
Farbentour
9
1
1
2.56
15.0%
Bavaria AI
8
2
3
2.25
13.3%
ithelps
8
1
1
3.62
13.3%
sunzinet
8
1
0
3.00
13.3%
Claneo was the broadest high-frequency result. It appeared seven times in ChatGPT, twice in Perplexity, six times in Claude and nine times in Copilot. Gemini didn’t recommend it.
Seowerk followed a different pattern: eight Gemini recommendations, seven from Perplexity and three from Claude, with none from ChatGPT or Copilot. Peak Ace received six Gemini recommendations, two from Perplexity and nine from Copilot.
Coinbound appeared only three times but reached three platforms. That distinction is useful: raw frequency measures how often a brand is selected, while platform reach measures how portable that visibility is across different systems.
Each AI platform formed its own shortlist
Platform
Recommendation rows
Unique agencies
Citation rows
Unique source URLs
ChatGPT
51
28
88
50
Gemini
50
10
32
7
Perplexity
60
25
103
56
Claude
45
17
62
38
Copilot
60
17
94
39
ChatGPT produced the broadest agency set
ChatGPT recommended 28 unique agencies across 51 slots. Its most frequent names were SNK and Claneo with seven appearances each, followed by Seokratie with five and RelioMedia with four. This was the least concentrated commercial recommendation pool in the test.
Gemini repeatedly returned a small group
Gemini filled 50 recommendation slots with only 10 unique agencies. pageworkers appeared 10 times, Farbentour nine times, Seowerk and sunzinet eight times each, and Peak Ace six times. Gemini also reused just seven normalized source URLs across 32 citation rows.
That concentration can create strong visibility for a selected group, but it also shows why a single-platform result can mislead. Three of Gemini’s four most frequently recommended agencies didn’t appear in any other platform’s recommendation pool during this run.
Perplexity returned the most citation rows
Perplexity produced 103 citation rows and 56 unique source URLs, the highest totals in the benchmark. Its leading agencies were ithelps with eight appearances, then Radyant, Seowerk and taismo with seven each.
Claude disclosed source reuse
Claude returned 45 agency recommendations and 62 citation rows. Its metadata said web search was run once for the batch and that results were reused where relevant. Kopp Online Marketing Consulting, Claneo and Bavaria AI each appeared six times.
This disclosure is important. We instructed every system to treat the 24 tasks independently, but a batch prompt can’t guarantee internal isolation. Claude was the only platform that explicitly described reuse across tasks.
Copilot completed every list but included placeholder citations
Copilot produced the maximum 60 recommendation rows. Peak Ace and Claneo each appeared nine times, Suxeedo eight times and morefire seven times.
Five Copilot citation rows pointed to example.com placeholder URLs. We retained them in the raw dataset because they are part of the observed output, but marked them as placeholders. They should not be treated as valid evidence.
What the source data tells us
The five systems returned 379 citation rows representing 186 normalized URLs. The most frequently cited domains were:
Domain
Citation rows
claneo.com
16
developers.google.com
15
masters-of-search.com
13
seowerk.de
12
seokratie.de
10
suxeedo.de
9
taismo.de
9
peakace.agency
9
farbentour.de
8
bavaria-ai.com
8
This is citation frequency inside the returned answers, not an authority score and not an independent endorsement of every page. The pattern still shows something practical: first-party service pages can strongly influence agency-selection answers, while technical prompts also pull from platform documentation and research.
A B2B company therefore needs both types of evidence. Its own website has to explain the service, audience, location, industries, process and proof clearly. Independent publications, profiles, comparisons and expert mentions then help other systems verify those claims.
Optis Digital’s baseline was zero
Optis Digital wasn’t named in any of the 120 answers. More specifically, it received zero recommendation slots across the 60 commercial agency-selection evaluations and zero mentions in the 60 informational evaluations.
That result is useful because it establishes a clean baseline. It doesn’t prove that Optis Digital is unsuitable for B2B GEO work. It shows that, during this test, the five systems didn’t retrieve or select enough evidence to include the brand for these unbranded buyer questions.
The gap is also visible in the winning domains. Several frequently recommended agencies had dedicated GEO pages that were repeatedly cited or used as evidence. Optis Digital now has a detailed Generative Engine Optimization service page, but one owned page is only part of the work. The brand needs repeated, consistent association with B2B GEO, SaaS, fintech, crypto, cybersecurity, Germany and DACH across both owned and independent sources.
What we’d do next to improve B2B AI visibility
1. Build category evidence around specific buying situations
Generic “GEO agency” positioning is too broad. The benchmark prompts reveal concrete situations that AI systems must resolve: a SaaS company entering DACH, a Munich software business, a fintech team that needs citation monitoring, or a crypto company looking for technical and authority support.
Optis Digital should publish pages and evidence that connect the company to those exact situations. The existing SaaS SEO service page can support this, but the relationships between SaaS SEO, GEO monitoring, international expansion and commercial outcomes should be explicit.
2. Publish data competitors can’t copy from generic documentation
This benchmark is one example. Future research can measure citation volatility, platform overlap, answer changes after technical fixes, AI referral landing pages and the difference between German and English prompts. Each study should expose its date, sample, method, raw results and limitations.
3. Earn independent category mentions
Owned service pages can be cited, but third-party evidence reduces the dependence on self-description. Relevant agency roundups, podcast appearances, expert commentary, conference profiles, partner pages and industry publications should describe the same specializations with consistent organization and founder details.
4. Track platform reach separately from total mentions
A campaign can increase recommendation frequency inside one platform without improving visibility elsewhere. Reporting should therefore separate total recommendation count, number of platforms, first-position count, average position, cited URLs, answer accuracy and AI referral traffic.
5. Repeat the fixed benchmark
AI answers change. A useful monitoring program keeps the core prompt set stable, records the platform and visible model label, runs from a known location and compares the new result with the original baseline. New prompts can be added, but they shouldn’t silently replace the benchmark used to measure progress.
Market focus: Germany, Bavaria, Munich, DACH and Europe, with SaaS, fintech, crypto and cybersecurity scenarios.
Location: Bayreuth, Bavaria, as shown or inferred by the platforms.
Branding: The prompts didn’t mention Optis Digital.
Recommendation rule: A company counted only when it appeared in the platform’s ordered recommendation array.
URL normalization: Tracking parameters and Markdown wrappers were removed for unique-URL counts.
Name normalization: Case variants and clear legal-name variants, such as Seowerk and Seowerk GmbH, were merged.
Raw-output handling: Format defects were repaired without changing substantive answers. The original defects are documented in the dataset.
Limitations
This is a reproducible snapshot, not a permanent ranking of GEO agencies. Each platform was tested once. Generative outputs can change with the model, mode, location, wording, personalization, retrieval index and date.
Three interfaces didn’t expose a precise underlying model name. We recorded the label shown rather than guessing. Perplexity returned an incorrect 2025 date in its metadata; we retained that raw value and normalized the test date to 2026. Claude added trailing non-JSON text, Gemini returned three unescaped quotation pairs, and ChatGPT and Gemini wrapped some URLs as Markdown links. Those issues were repaired during normalization.
We recorded citations as the platforms returned them. We didn’t independently validate every claim made by every cited page. Five Copilot citations used explicit example.com placeholders and are marked accordingly. Agency inclusion should be read as observed AI visibility, not as an Optis Digital endorsement or a quality ranking.
Frequently asked questions
Which agency had the strongest visibility in this B2B GEO benchmark?
Claneo had the strongest combined result, with 24 recommendations across four of the five platforms. It appeared in 40% of the 60 commercial prompt-platform evaluations.
Did any GEO agency appear in all five AI platforms?
No. Claneo reached four platforms. Seowerk, Peak Ace and Coinbound each reached three. No agency appeared in ChatGPT, Gemini, Perplexity, Claude and Copilot during the same benchmark run.
Which AI platform returned the most sources?
Perplexity returned the most citation rows: 103 citations representing 56 unique normalized source URLs. ChatGPT followed with 88 citation rows and 50 unique source URLs.
Was Optis Digital recommended?
No. Optis Digital received zero recommendations across the 60 commercial evaluations and wasn’t mentioned in the 60 informational evaluations. This result is the baseline against which future GEO work can be measured.
Were these 120 different prompts?
No. The study used 24 unique prompts across five AI platforms, producing 120 prompt-platform evaluations.
Can this benchmark prove which GEO agency is best?
No. It measures which agencies the five systems selected for a defined prompt set on one date. It doesn’t independently assess service quality, client results, pricing or fit.
With years of experience navigating the ever-evolving crypto landscape, Eugen knows exactly how to make content shine in Google’s eyes—without breaking the algorithm. With experience working as an SEO specialist in real fast-growing crypto companies, along with training in crypto trading, Google Ads Search Certification, and Google Analytics Individual Qualification, he is a master of SEO in the crypto world, blending AI-powered strategies with deep industry knowledge. From ChatGPT to blockchain trends, he knows how to make content rank, engage, and convert.