Is generative engine optimization worth it? What 3 measurements showed (2026)

geo-ai-seo7 min leestijd

Search for generative engine optimization and you find two kinds of texts: sales pages calling it indispensable, and sceptical pieces calling it hype. Both camps share a problem: little measurement data. We measure AI search visibility for a living, including for our own site. This article answers the sceptical questions with what 3 measurements showed, including the numbers that argue against us.

The question as sceptics ask it

The sceptical questions are fair and concrete. Is it worth the cost, or are you paying for something you cannot verify? What are the downsides and risks? Is it hype that will blow over, or worse: are providers selling promises nobody can keep? And can you not simply do it yourself?

Our position up front, so you know who is writing: we sell measurements and improvement work in this field. That is exactly why we publish the numbers that partly prove the sceptics right.

What we measured

Measurement set: 5 questions about our own field, 3 measurements between 4 and 6 September 2026, on 4 platforms (configuration in the framework below). The outcomes:

1. ChatGPT mostly did not search. For 4 of the 5 questions ChatGPT (gpt-4o via the Responses API, with search tool available) performed no web search, in all 3 measurements. The model answered from its own knowledge. This is the strongest sceptical data point we have: for those 4 questions no content route existed at that moment. No optimisation would have put us in those answers. Whoever promises you a mention for such questions is selling something they cannot steer.

2. The other platforms searched extensively. Claude (claude-sonnet-4-6 with web search) fired 13 to 15 search queries per measurement and consulted 93 to 120 sources, of which 38 to 46 were cited. Perplexity (sonar-pro) consulted 96 to 99 sources per measurement and quoted 52 to 57 of them in the text. Where searching happens, real pages get read and cited; there the content route very much exists.

3. We were not among them ourselves. Across the 3 measurements ChatGPT and Claude recorded 32 unique search queries. In 0 of them a page of our own site appeared among the consulted sources; 184 other domains did. On the question about a digital strategy agency in Utrecht, Claude consulted sortlist.nl and digitalinside.nl, among others. Directories, comparison articles and third-party city pages weigh heavily in what the platforms read.

4. We were rejected nowhere. There were 0 cases where one of our pages was read but not cited. Our gap sat entirely in being found, not in being judged and rejected. To illustrate how concrete that gets: on 9 September 2026, 2 of our own knowledge base pages turned out not to be in the Google index. That is not an abstract AI problem; it is an indexing problem with a name, a URL and a fix.

The honest summary of 3 measurements: part of the playing field is closed (you cannot influence not-searching), part of it is simply open, measurable work (indexing, content, mentions in the places platforms consult).

What this means for doing it yourself or outsourcing

What you can do yourself. Check whether your pages are indexed (Google Search Console is free), get structured data and robots.txt in order, and write content that answers your customer's question. That is the foundation, and it needs no agency.

Where it gets harder. Knowing which questions your customers ask AI, measuring whether you are mentioned across platforms, and above all: determining per non-mention which category the cause falls into. Without that distinction you do not know whether your content work makes sense or whether you are optimising for a question where the model does not search anyway.

What an agency cannot promise. A mention in ChatGPT. Our own data shows why: for 4 of the 5 questions the model did not search, and what a model answers from its own knowledge cannot be steered with content. An honest agency says so out loud and shows it in the measurement.

What an agency can deliver. A baseline measurement, the cause category per gap, focused work on the open category, and repeat measurements on exactly the same question set so you can see the effect. Cost and scope differ per situation; the measurement makes them bounded and verifiable instead of open-ended.

How to check whether an agency measures or merely promises

Five control questions, all askable in one conversation:

  1. Ask for the measurement configuration. Which models, which APIs, how many questions, how many repetitions, which dates. Whoever cannot provide that does not measure.
  2. Ask how non-mentions are explained. The answer must distinguish between did-not-search, not-found and read-but-not-cited. Whoever attributes everything to "more content" does not have that distinction.
  3. Ask what is not possible. A provider who never says that not-searching cannot be forced is selling beyond their data.
  4. Ask for the question set. Repeat measurements are only comparable on a fixed, pinned set. If the set changes per measurement, every trend is a story.
  5. Ask for a result that disappointed. Whoever has none either does not measure or publishes selectively.

Frequently asked questions

Is generative engine optimization worth the cost? That depends on where your absence comes from, and that is measurable. If the gap sits in not being found (as with us: 0 of 32 search queries with an own page among the sources), the work is concrete and verifiable. If it sits in questions where the model does not search, there is no content route and you should know that up front.

What are the risks? The biggest risk is paying for promises that cannot be verified. The remedy is not abandoning the whole field but measuring: a baseline measurement costs little and makes every follow-up claim testable.

Is it hype? The sales talk around it contains hype, and our own ChatGPT numbers show part of the promise is impossible. The measurable part underneath (being found by systems that demonstrably search and cite) is not hype but plain work.

Can you do it yourself? The foundation, yes. The measuring and the cause distinction per non-mention require tooling; how Lens compares to other measurement tools on publicly documented properties is in the AI visibility tools comparison (2026). In any case, start yourself with the free check whether your pages are indexed; that was our own first finding.

The measurement method

All numbers in this article come from this configuration and do not generalise beyond it:

  • Measurement set: 5 fixed questions about our own field (AI search visibility and agency selection), pinned in advance and identical in every measurement.
  • Repetitions: 3 measurements, between 4 and 6 September 2026.
  • Platforms and configuration: ChatGPT via the Responses API with search tool available (measured model gpt-4o-2024-08-06), Claude with web search (claude-sonnet-4-6), Perplexity (sonar-pro), Google AI Overviews via real Google search results.
  • Recorded per answer: whether the platform searched, the fired search queries, the consulted and cited sources, and the mentions.
  • Limitations: this is one site, one question set and one measurement window. We do not measure or claim training effects of models. Outcomes for other sites and questions can differ; that is exactly why your own baseline measurement is the starting point.

Want to see what this looks like for your site: request a free Snapshot or first read what AI search visibility is.

How does your business score?

Request a free AI Visibility Snapshot: 1 page, no strings attached.

Request a Snapshot →