← The library
Format
Research Essay
Reading time
42 min
Reading level
Advanced
Published
29 July 2026
Topics
Artificial Intelligence
Technology Strategy
Platform Strategy
Digital Commerce
Customer Experience

Research Essay 02

AI Visibility. What Can Be Measured and What Is Merely Sold

An evidence-based reading of AI Search, GEO and Local AI Visibility for multi-location brands

AI-mediated discovery is a real, probabilistic management problem, not a stable ranking system. What can be measured today, what a brand can actually influence, and where uncertainty is being sold as certainty are three different questions.

Editorial note

This first publication edition names technical platforms, research work and institutional sources where their identity is necessary for a claim to be traced. Concrete examples of commercial measurement and optimisation products, on the other hand, are anonymised. Their originals, wording and access dates remain documented in the private evidence edition. Anonymisation is not meant to place a critique beyond reach. It is meant to limit it. Absence of public documentation is not read as absence of internal capability, and an over-large claim is not treated as evidence of poor product quality or bad faith. The essay tests what a publicly accessible source can carry and where its reach ends. This edition therefore contains neither a vendor league table nor the claim that a single tool can map the market in full. It develops a measurement, procurement and operating model for brands that need to act while part of the system remains unobservable.

The market is selling a ranking

A multi-location brand across the DACH region has a new item on its agenda.

Teams in Germany, Austria and Switzerland want to know whether the brand is visible enough inside standalone AI assistants and the generative surfaces bolted onto established search products. A vendor could show how often the brand is mentioned or cited. A competitive comparison could rank who appears more. The result could be tracked as a weekly score. Neat, tidy, and quietly wrong in several places at once.

This is not a reported client case. It is a familiar decision situation. Something hard to see is handed a number, the number is handed a rank, and the rank finds itself sitting next to a budget question that nobody asked out loud.

Should the brand now invest?

A ranking is a seductive answer. It brings order to a topic that pretends to be technology, marketing, market research and attribution at the same time. Unfortunately it collapses questions that live in different populations.

How many people use generative systems for search at all. On which surfaces do they meet AI answers. What does a synthetic prompt panel actually see. How often does a click follow. And does that contact make its way into a visit, a lead, a booking or revenue.

These questions are connected. They are not interchangeable.

AI-mediated discovery is relevant enough, for specific intents, segments and product surfaces, to be observed systematically. Nothing about that observation licenses a universal AI-search market share or a wholesale budget shift.

The market is selling transitions

The story of the new channel is plausible in outline. People use generative systems. Some of them look for information, products or local services there. Brands appear in the answers, occasionally with a link. Some of those contacts influence a later decision. From a distance, that resembles a customer journey.

GenAI usage
→ AI search
→ visibility
→ real exposure
→ engagement
→ business result

Economic value is imagined at the end of the chain. It is often sold at the beginning.

General ChatGPT or Gemini usage may mean writing, translation, coding, image generation or entertainment. Information retrieval inside a chatbot is not proof that classical web search has been replaced. A synthetic answer shows what a defined test system produced under defined conditions. It does not show how many real people asked the same question or saw the same reply. A visible source citation is not a click. A click is not an incremental business effect.

None of these limits makes AI search unimportant. They only decide how far a number is allowed to travel.

That is the current tension. Discovery is genuinely changing. Generative answers are being folded into established search products, standalone AI services are being used for information tasks, and familiar click paths are shifting. At the same time, there is no single numerator and denominator that turn all of this into a market share. The market therefore sells not only monitoring or advice. It sells the confidence that the transitions have already been measured.

Relevance, for a brand, is not a disguised market-size question. A behaviour can matter strategically long before it accounts for the largest share of a journey. It matters when it touches a valuable audience, occurs at a critical decision point, or produces errors that can be found and fixed at reasonable cost. A quickly growing behaviour can equally mean little to a specific business if it mainly concerns other tasks, regions or categories.

Relevance is therefore the value of a better decision under uncertainty. What would it cost to ignore the signal. What does it cost to check your own situation. Which actions have a defensible payoff regardless, and which would only ever be funded by a new ranking. These are the questions that do not require an invented total. They require a clear line between observation and action.

Nineteen and eighty

The strongest direct DACH anchor for local service discovery shows how quickly two respectable numbers can be conscripted into a story larger than either of them.

For the KMU Digital Pulse 2025, the Lucerne University of Applied Sciences and localsearch surveyed 1,660 people from the three main language regions of Switzerland online. Fieldwork ran from 23 June to 2 July 2025. A Ticino oversample was reweighted for the overall analysis. Nineteen per cent reported that, in the past twelve months, they had used generative AI to search for and inform themselves about services from small and medium-sized businesses. Eighty per cent named search engines as their first port of call when a provider was still unknown.1

The finding is real. Almost a fifth of a weighted Swiss sample recalled a concrete AI-mediated information use in a local service context. AI was not a future idea in that journey.

The two numbers do not, however, produce a new market share. Nineteen per cent describes people who did something at least once in a year. Eighty per cent comes from a differently worded question about first port of call. These are not complementary shares of the same population. The survey observed no actual queries, no delivered answers, no clicks and no purchases.

A defensible reading stays deliberately narrow. In Switzerland, generative AI is already part of local service journeys. Classical search engines remain, by a comfortable margin, the more common first port of call. That is all the finding carries. It shows no displacement rate. It does not say what share of local demand for a specific multi-location brand originates in AI. It does not confirm that people in Germany, Austria and Switzerland use the same systems for the same categories, intents and location types.

For attention, this is enough. For a permanent budget shift, it is not.

Five things under one name

"AI Search" describes very different measurement objects in market reports and dashboards. A small table helps more than a grand definition.

One term, many systems

ObjectTypical unitWhat it answersWhat does not follow
general GenAI adoptionshare of people over a periodHave people used generative systems?share of AI-search queries
active AI information or service searchpeople or reported tasksWas AI used for a defined search task?market share of all searches
synthetic AI outputanswer, prompt run, panel indexWhat appeared under defined test conditions?real population exposure
AI referralsession or visit with source signalWhich measurable clicks came from an AI surface?total, including non-clickable, effect
business outcomecall, route, lead, booking, purchase, revenueWhich result was observed or attributed?incremental effect without a comparison

These are not rival versions of the same truth. They are sensors for different slices of a journey. A survey shows what people remember and which port of call they name. A browser panel observes visits and clicks. Web analytics captures referrals when a source signal is transmitted. A prompt tracker collects answers under a fixed configuration. First-party data can record calls, routes, leads, bookings and purchases. None of these methods sees the full chain. Which is precisely why each of them can be useful.

Work does not begin with the search for the one right number. It begins with a decision about which question is being answered. A CMO may want to know whether a new discovery behaviour is emerging. A local team is looking for wrong branch information. Analytics wants to trace the transition to an observable outcome. Procurement wants to understand what a monitoring product actually measures. A single visibility score can compress those tasks. It cannot perform them for one another.

The screen is not yet the market

The most consequential confusion is the one between what a test produces and what a population actually sees.

A prompt panel can repeat the same local question on defined platforms. It can document whether a brand is mentioned, recommended or cited. For monitoring, diagnosis and experiments, this is valuable. What the panel does not know is the real frequency of that question. It does not know which wording people use, which personal or local context is live, or whether the same answer is served to them at all.

Synthetic output is an observation under defined conditions. Real exposure requires an additional bridge.

Inside established search products the picture is no cleaner. A person can meet an AI Overview without ever deliberately choosing an AI service. Someone else can use a chatbot for information without producing a single web referral. "AI Search" can mean active choice, passive exposure, synthetic observation, or attributed traffic. Self-report and telemetry show the same problem from opposite ends. The Reuters Institute asked people in six countries about their use of generative systems and their behaviour around AI-generated search answers. Pew reconstructed actual US browsing by a panel and compared Google visits with and without an AI Overview.2

The findings sit far apart and refuse to combine into a single click-through rate. Reuters measures recalled frequency at the person level across six countries and several AI-search contexts. Pew counts observed visits in a US panel on a specific Google surface. Unit, country, period and method differ. That does not disqualify either. Self-report can capture perception, trust and non-clickable use that never touches web analytics. Telemetry observes visits and click events that people remember imprecisely. Together they show more than either alone, so long as their denominators stay visible.

The sensor is not the market

Economic change is more real than the ranking

Methodological restraint should not slide into the comforting claim that nothing economic is happening.

A randomised US field experiment found that, for queries which actually triggered a Google AI Overview, outbound organic clicks fell. A separate quasi-experimental study observed a material drop in informational Wikipedia use after AI Overview exposure. Both concern specific US settings, one of them close to publishing. Neither is a DACH study, a local-search experiment or a revenue analysis. Together they strengthen a narrower claim. Generative answer surfaces can shift attention and click paths.3

For publishers whose business hangs directly on pageviews, this effect looks one way. For a multi-location brand it can look quite different. A local customer may read an opening time, search the brand directly later, start a route, phone the branch or walk in. An AI answer can replace a click without preventing a purchase. It can also create attention without sending a visible referral.

Referral data remains useful when its reach is not overstretched. A 2026 Marketing Science analysis of 973 e-commerce sites found economically measurable but low-volume organic ChatGPT referrals. Conversion rate and revenue per session in the main analysis sat below most established channels. The study is descriptive, uses proprietary data and offers no DACH breakdown. It therefore neither proves a general quality disadvantage of AI traffic nor a causal channel effect.4

The more useful business question is not whether "AI traffic" converts better in the abstract. It is what volume, what user selection, what cost and what incremental contribution appear in the concrete journey.

A brand is not a market

The multi-location brand does not need the total AI-search market share to act. It needs three things. Where do its customers show up, in which intents, and where does an early observation change a decision that would otherwise be made blind. Early indicators are enough to start. What is disciplined here is not the vocabulary but the boundary between observation and promise.

A first budget rule follows. Fund only what has a defensible payoff regardless of the AI-search ranking, plus a modest, honestly labelled test that can be shut down without face-saving. Everything else waits for evidence that a specific journey is real, measurable and worth changing.

Before the sensor question, one prior question deserves an answer. Did the person even perform a search in the sense that the dashboard assumes.

General usage of a generative system is not the same as search intent. Someone drafting an email, brainstorming a birthday present or asking for a recipe is not shopping. If the tool volunteers a nearby restaurant, that is closer to a suggestion than a query. Counting it as an AI-search event inflates the numerator without touching the denominator. Any measurement stack that cannot distinguish task type from moment of exposure will overstate its own subject.

A map, not a blueprint

Generative product surfaces differ from one another and from themselves over time. ChatGPT Search may be triggered automatically or manually, rewrites queries, shows citations, and can act on approximate location or a memory profile.5 Google separates AI Overviews from AI Mode, uses conditional serving and a technique publicly described as query fan-out.6 Microsoft distinguishes consumer Copilot from enterprise Copilot with different web-search policies.7 Perplexity offers modes and an API with distinct behaviour.8 Anthropic exposes web search as an activatable, policy-controlled tool state.9 Gemini shows sources in-app and a structured grounding metadata channel via its API.10

None of these documentations publishes a full retrieval, ranking or context-allocation logic. Each names a surface, a tool and a policy state. Together they draw a map. They do not deliver a blueprint. A vendor claim that a single tool captures "all AI answers" is either careless language or a sales instinct with poor sense of humour.

The prompt is not alone

An AI answer is the last visible stage of a pipeline. Between prompt and reply sit query rewriting, retrieval from an index or a live web call, ranking, filtering, grounding and, on some surfaces, memory and location conditioning. OAI-SearchBot, Google-Extended, PerplexityBot and comparable agents are technical preconditions for a page to be reachable at all, not evidence that it will be reached.11

One brand, many local realities

A source link in a chat interface is a compact instruction to look further. It is not a peer-reviewed endorsement of the linked page, nor a guarantee that the sentence next to it is supported by that page. Research on generative search engines and on source attribution documents both stated citations and the frequency with which those citations fail to support the accompanying claim.12 A dashboard that treats citations as trust weights inherits a problem the underlying systems have not solved.

Place is part of the answer

Locality is not a filter that a brand switches on. Depending on surface and setting, the system uses IP-based approximate location, an operating-system location signal, a stored profile, a memory entry, or a Business Profile record ingested elsewhere.13 A brand appears not because a canonical URL was crawled but because a set of identity, review, address and category signals converged closely enough to the question in front of the model.14

The practical consequence is that "the brand" is technically several things at once. Corporate homepage, product page, structured data on those pages, Business Profile per branch, product feed, review corpus, third-party listings. Each of those has its own maintenance discipline, and any measurement plan that assumes they are one object will find itself measuring the least well-kept of them by accident.

Object before metric

Any credible measurement design begins with an observed object. What is the population, the surface, the query set, the time window, the language. A dashboard that reports a "brand visibility score" without those attributes has not measured a brand. It has measured the effort of its own vendor.

The right first question is therefore not "what is our score" but "what is the sensor actually looking at". A prompt tracker sees the output of its own configuration. A survey sees remembered behaviour. A panel sees observed sessions from a limited set of devices. Web analytics sees what an AI surface chose to transmit as a referrer. First-party data sees the entry to the physical or transactional edge of the business. None of them measures the market.

A metric needs a task first

Numbers without tasks tend to grow decorative. A useful indicator answers a specific question that a specific role will act on. Ranked appearance in a prompt panel across a curated question set is a diagnostic for the content and profile inventory of a category, useful for the local team and the SEO lead. Referral traffic from AI sources is a channel-quality measurement for analytics. Google Search Console's dedicated generative AI performance report gives URL impressions on named Google surfaces, which is a first-party first step, not a cross-surface exposure measurement.21 Standard traffic-source dimensions in analytics carry engagement, subject to well-known non-click and app gaps.22

What the sensor actually sees

Two failure modes account for most inflated numbers. The prompt set is either too polished or too random, and the denominator has been quietly dropped.

An in-house question inventory of a few hundred prompts, hand-written by the brand team, tends to over-represent branded queries and under-represent the messy long-tail that a real customer produces. A machine-generated question set from a language model contains the biases of that model. Neither is disqualifying. Both need to be labelled, versioned and periodically compared against a small externally sampled reference set, ideally with a real research counterpart. Guidance on composite indicators from OECD and Joint Research Centre and disclosure standards from AAPOR set the general expectation for how such measurement should be documented.1520

The prompt universe nobody fully knows

Generative outputs vary between runs even at low temperature. Research on stochastic evaluation of large language models has produced usable recipes for how many prompts and runs a stable estimate needs.16 Complementary work suggests that some apparent prompt sensitivity is an artefact of evaluation method rather than an intrinsic property of models.17 Both bodies of work are settings-specific. Neither hands the brand a universal number of prompts or runs. Both make one thing clear. A single-run screenshot is decorative. A repeated, versioned measurement stands a chance.

Public first-party accounts of how people use ChatGPT describe broad task categories but do not equip an external brand with a market and location-specific prompt frame.18 Any commercial claim to know the full prompt universe of a category ought to be received with the polite scepticism reserved for weather forecasts more than a week out.

Preconditions are not ranking factors

More data does not fix every error

Sample size fixes noise. It does not fix bias. A large panel drawn only from desktop browsers still misses mobile assistants. A sizeable prompt corpus still misses the natural distribution of questions. A long time series still cannot answer a question about a period before the sensor existed.

Two anonymised commercial illustrations from the essay's documented claim sample sit here rather than in the marketing section, because they are diagnostic. In one, a monitoring product reported user behaviour without documenting a validation against observed purchases. In another, a supplier described its product with limiting language that made the offer honest but the score less spectacular.19 Both cases carry the same lesson. What the label promises has to survive the smallest audit request in procurement.

The denominator has a vote

An AI answer that appears in a synthetic panel does not automatically reach a person. A search behaviour that a survey reports is not automatically visible in a panel. A click that arrives in web analytics is not the total effect of an AI mention. Institutional counter-factual methodology treats a comparison group as a precondition for any causal claim.23 A visibility score without a comparison group is a description at best.

Four evidence levels

It helps to keep four levels of evidence separate. Presence, whether a brand appears in an output at all. Prominence, whether the appearance is early, cited and correctly attributed. Exposure, whether real people meet that output in meaningful numbers. Effect, whether a measurable change in behaviour or outcome follows. Each level has its own sensor. Each has its own denominator. A programme that treats them as a single "visibility" metric will optimise the cheapest one.

A minimum viable AI-visibility report therefore carries five pieces at once. Object and unit. Sensor and denominator. A short list of decisions the report is meant to inform. A note of what the report cannot conclude. And a versioned prompt or query base with an audit trail.24

Three answers under one name

"Local" in an AI answer can mean at least three things. The system inferred a location from a network signal. The user stated one in the prompt. The user's account carries a stored profile. The three routes are not equivalent, and the same brand may be handled well by one and clumsily by another.

Where is "here"?

Signals from surfaces such as Google Search, ChatGPT and Claude are documented as potentially, and unevenly, using location.13 None of the platforms publishes a constant weighting of that signal in every answer. That is a design choice, not a lapse. It leaves the brand with a modelling problem rather than a reading task.

Which branch is meant?

The GEO promise machine

Locality can travel through several data paths at once. Business Profiles carry hours, address, categories and reviews.14 Product feeds carry stock, price and shipping. Structured data on canonical pages carries organisational and product identifiers. Third-party sites carry citations and reviews. A brand that maintains only one of those channels well will find that "the brand" behaves inconsistently across surfaces without any user-facing reason.

Exhaustive, distributed or targeted?

Measurement across a large branch portfolio faces the sampling question every operator eventually has to name. Test every branch shallowly, sample a representative subset deeply, or target only branches with known problems. Each choice carries a bias and a cost. What matters is the note that follows the number. This is what was tested, this is what it can say, and this is where an average will mislead if used without a way back to a specific branch.

From answer to local outcome

The outcome side of a local journey depends on identifiers the brand controls. A store locator with clean structured data, a per-branch landing page with correct primary category, working phone and route buttons on the profile, and consistent review response over time. Without those, an AI answer that mentions the brand still meets a broken hand-off. Advanced tooling on the sensor side cannot compensate for a broken handrail on the outcome side.

DACH is not a measurement condition

Filing three countries under one heading is administratively convenient and analytically expensive. Languages, dialects, review platforms, local media, category conventions and legal expectations differ. Any panel report that averages across DACH without a country breakdown has already thrown away the information most useful for a brand with country-specific field teams.

A usable portfolio truth

The most useful portfolio view is not a single visibility number but a small matrix. Country, city size band, primary category, evidence level. Each cell carries a modest observation and a decision the observation would trigger. That is enough. It saves the brand from the temptation of a single beautiful score that no one can act on.

The error is usually in the reasoning

Most controversial GEO claims are not lies. They are conclusions drawn from a base that does not support them. A vendor tests one site or a small handful, finds a promising pattern, and describes the pattern as a general rule. The pattern may be real for that site. The generalisation is a step too far, and the loudest step usually earns the largest fee.

Where claims become too large

Three moves account for most oversized claims. A finding from a single case is presented as a rule. A synthetic panel is described as market truth. An intermediate metric such as citation or presence is described as an effect on decisions or revenue. Each of these moves has a legitimate lesser form. A case can illustrate. A panel can indicate. An intermediate metric can be tracked. Claim discipline lives in the words between illustration and rule.

Tools are a stack, not a league

What a real basis includes

A serious optimisation claim needs a base, a comparison group, an observation window that survived the model updates it lived through, and a note on the confounders that were controlled and the ones that were not. NIST AI RMF and standard disclosure guidance describe what "measured" ought to mean in a discipline that would like to be one.15

From a basis to a plausible contribution

Contribution is not causation. A brand that improves structured data, canonical URLs and Business Profile hygiene will find that some appearances rise, some fall, and some are unchanged. That is a plausible contribution, not a demonstrated causal effect. Both are useful. Only one licenses a claim about "increasing your AI visibility by X per cent". The other supports the more honest statement that basic hygiene is being maintained and its effect is being observed.

The known GEO finding at the right size

Recent research on generative-engine optimisation shows that certain content characteristics are associated with a higher probability of citation in generative search products.15 The correlations are worth taking seriously. They are also settings-specific and change with model updates that are neither pre-announced nor documented for the outside world. A brand that optimises exclusively for a specific model at a specific time is optimising a target that moves.

A portfolio tests differently than a page

An optimisation that helps one page can hurt another. Aggregate optimisation across a large branch or product portfolio requires a study design that keeps track of interactions. A modest cluster-randomised test with a well-labelled control group is worth more than a large uncontrolled rollout, because it can be defended after the model updates the world underneath it.

Brands must act before perfect evidence

None of this argues for waiting. It argues for acting with the right size of promise attached. A brand that maintains its canonical identity, its structured data, its Business Profiles, its product feeds and its review discipline is not chasing a ranking. It is keeping the roads open, whichever surface the traveller uses.

The decision of the multi-location brand

The brand at the top of this essay does not need to buy a visibility monitor to justify the work above. It needs a decision about scope, a small versioned prompt base, a first-party measurement of AI referrals from Search Console and analytics, and a per-branch check on Business Profile hygiene. That is enough to start. Everything more expensive should wait for a question the smaller programme could not answer.

Four fields, nothing more

An honest procurement conversation about a monitoring product fits on one page. What does the product measure. What is the population it measures against. How is it validated over time. And what decision, taken by which role, does the output support. If any of these four fields is missing on the vendor side, the answer is not a bigger meeting. It is a smaller purchase.

From observation to promise

Marketing rewards clean promises. Measurement rewards visible caveats. The tension is real, and it is manageable. A promise translated from an observation should mention the object, the sensor, the population, and the decision it is meant to inform. That reads longer than a slogan. It is also the reason procurement will approve it a second time.

A workable operating model

Three miniatures from the documented claim sample

One vendor reports a national visibility rank without a per-branch drill-down. The rank moves attractively over time. It does not answer a question any operator in the network can act on. Another vendor tracks citation frequency across four models but treats them as an average score. The average conceals divergences that are the actual news of the report. A third vendor sells a "GEO optimisation" package that leaves it unclear which changes were made to which pages when. Each of these is anonymised because the point is not the vendor but the shape of the claim. In every case the fix is smaller than the complaint. Publish the object, the unit and the change log.

The national score before the branch

A rank that averages across a country hides everything the local team could use. A brand with 250 branches whose national visibility rank is stable can have a quarter of its network invisible in local answers, and the report will look fine. Any measurement stack that cannot descend from the country score to the individual branch on demand is a poster, not an instrument.

How short is marketing allowed to be

Short. Not fictional. A single sentence can carry a truthful claim if the sentence names an object and a sensor. "Cited in twelve per cent of tested prompts across four models between May and July, from a versioned base of 800 branded and unbranded questions in DACH." Marketing colleagues will find this ugly. Procurement will find it welcome. Legal will read it twice and not intervene. That is the trade.

Uncertainty is manageable

What the claim must survive in procurement

Procurement is not the enemy of a good measurement product. Procurement is the discipline that asks whether the product does what the label promises. The moment a vendor is unwilling to sit through five questions about population, sensor, denominator, validation and change log, the meeting has ended the vendor a favour by keeping it short.

What should work after purchase

Tools solve a task inside a system. They do not build the system. A useful AI-visibility stack for a multi-location brand consists of a small handful of components with clear responsibilities, none of them heroic. A prompt tracker with a versioned question base and a labelled model coverage. Search Console and analytics for first-party AI referrals. Business Profile insights for local behaviour. A per-branch content and hygiene routine that a real team owns. That is a stack. Adding a tenth tool without deciding which decision it improves is a hobby.

The stack emerges from the work

Which tool sits where depends on the work. A team that has never held a versioned prompt base does not need the most advanced tracker on the market. It needs the discipline of maintaining a base at all. A team already running clean referral analytics does not need a second referral tool. It needs a proper hand-off from AI observation to editorial and profile work.

Five judgements that must not disappear into a score

Score-driven procurement makes it easy to lose five judgements. Which population is measured. Which surface is in scope. Which local reality is respected. Which change is being controlled for. And which role acts on the number. A one-line score can hide any of these. A short report should make each visible in a sentence.

Local starts before the country filter

A tool that only offers a country filter for local questions is not a local tool. Local is a city, a district, a branch, a category and an intent. A meaningful local view sits at least at city and category level and can descend to the branch when required. Country is a summary, not a location.

Data must survive the contract

A dataset that cannot be exported at reasonable cost is a lock-in, whatever the roadmap says. Contracts should specify export formats, retention, and the ability to move a versioned prompt base to another vendor without losing the history that made the measurement meaningful. This is not a legal decoration. It is what allows a brand to change tools without losing the argument.

The price needs a unit first, then a context

"Per user" and "per prompt" and "per market" price the same product very differently. A vendor that will not name the unit is not ready to sell. A vendor that names the unit but not the context of a realistic annual usage has not thought about the price the buyer will actually pay.

Why the research did not turn into a league table

The temptation to publish a vendor league table is strong. The research does not support one. A responsible ranking would need a shared measurement standard, an independent audit, and a stable comparison window that survives a year of model updates. None of these exists in a form that would let a public ranking mean what it appears to mean. A shorter ranking would mislead. A longer one would already be out of date.

The RFP as a test of clarity

A useful request for proposal on this topic is short. State the four fields above. State the two or three decisions the tool must support. State the export requirements. Ask for a joint pilot on a small, real portfolio rather than parallel sales demonstrations on curated material. A vendor who cannot join a pilot on those terms is telling you something.

A joint pilot beats parallel demos

Sales demonstrations are auditions. A joint pilot is a fitting. Pilots reveal how a vendor handles change requests, how transparent the measurement documentation actually is when running against a live portfolio, and how support behaves at three in the afternoon on a Friday. All three matter more than the polish of a demo.

Yes, procurement is allowed to pick a winner

Marketing does not have to hold the pen on this decision. Procurement, with marketing and analytics reviewing evidence, is a defensible constellation. The role that will live with the contract for three years is entitled to the strongest voice at the moment it is signed.

The real work starts after the contract

A tool that has been bought and not integrated is worse than no tool. The integration work is not glamorous. Version the prompt base. Wire the referral streams. Write the one-page monthly report the brand will actually read. Book the quarterly review that will decide what to keep. Nothing here is difficult on paper. It is only difficult in calendars.

Decision first, metric second

An operating model for AI visibility begins with the smallest useful list of decisions and the roles that will make them. A CMO decides whether the topic warrants strategic attention. A local operations lead decides which branches receive priority intervention. Analytics decides whether an observed change deserves an experiment. Procurement decides whether a vendor stays. If the operating model cannot name these roles, no dashboard will save it.

Thirty, ninety, one hundred and eighty

A calendar helps only when it is honest about being a convention. Thirty days for setup and a first read. Ninety days for a first internal review with real portfolio data. One hundred and eighty days for the first decision on whether the programme continues, changes shape, or closes. These gates are editorial, not industry standards.65 The real gates come from data state, portfolio, measurement design, organisation and decision.

Total cost of ownership, honestly named

Tool cost is the smallest line. Real cost includes an owner for the prompt base, an editor for the monthly report, engineering time for the referral wiring, a per-branch hygiene routine, a periodic external reference sample, and the review time of three roles at least. A TCO schema that does not name these components is a wish list, not a budget.66

A report worth reading

Whatever the stack, the monthly output belongs on one page. Object and unit. Sensor and denominator. What changed. What the change appears to license as a decision. What the report does not conclude. Two or three per-branch call-outs. Nothing more. If a page cannot carry a decision, the additional pages will not either.

When to close the programme

A closing rule protects a programme from becoming furniture. If three consecutive quarterly reviews cannot name a decision the programme informed, the programme is a subscription. Rename it, downsize it, or cancel it. Nothing dishonours a good measurement discipline more than keeping a bad one alive to save face.

Uncertainty is manageable

AI-mediated discovery is a real, probabilistic management problem. It is not a stable ranking system. What can be measured today, what a brand can actually influence, and where uncertainty is being sold as certainty are three different questions. Answering them honestly makes the topic smaller, cheaper, and considerably more useful.

The multi-location brand does not need to know the exact share of local demand that arrives via generative surfaces. It needs to keep its canonical identity clean, its Business Profiles healthy, its structured data current, and its referral analytics wired to the surfaces that actually transmit a source. It needs a small versioned prompt base, an owner who reads it, and a monthly page that a real person would read on a train. Everything else is optional.

The market will continue to sell a ranking. The brand does not have to buy one. What it needs, and what a discipline of measurement can give it, is a defensible line between what has been observed, what can be promised, and what should wait. A line, in the end, is enough.

Apparatus

Notes

1 Thomas Wozniak, Michael Boenigk, Marcel Niederberger and Nadine Stutz, "KMU Digital Pulse 2025", Lucerne University of Applied Sciences and localsearch, 2025, esp. pp. 8, 21 to 23 and 34 to 36, Whitepaper PDF, accessed 28 July 2026. Online survey, 1,660 respondents from the three main Swiss language regions, fieldwork 23 June to 2 July 2025, Ticino oversample reweighted. The 19 per cent figure is reported use within twelve months. The 80 per cent figure comes from a differently worded question about first port of call. Neither figure is a query share or a local purchase share.

2 Felix Simon, Rasmus Kleis Nielsen and Richard Fletcher, "Generative AI and news report 2025", Reuters Institute for the Study of Journalism, 7 October 2025, Reuters Institute. Athena Chapekis and Anna Lieb, "Google users are less likely to click on links when an AI summary appears in the results", Pew Research Center, 22 July 2025, Pew Research Center, both accessed 28 July 2026. Different units, not combined into a shared click-through rate.

3 Nikhil Agarwal and Rishabh Sen, "The Impact of Google AI Overviews on Publisher Traffic and User Experience", working paper revised 8 July 2026, SSRN. Mohammad Khosravi and Hema Yoganarasimhan, "Impact of AI Search Summaries on Website Traffic", preprint version 4 of 12 May 2026, arXiv, both accessed 28 July 2026. Randomised and quasi-experimental directional evidence for specific informational settings. Not a DACH, local or revenue estimate.

4 Maximilian Kaiser and Christian Schulze, "ChatGPT Referrals to E-Commerce Websites", Marketing Science 45(4), 2026, pp. 699 to 715, published online 21 April 2026, INFORMS, accessed 28 July 2026. Peer-reviewed descriptive analysis of 973 e-commerce sites using proprietary data. No isolated channel effect. No separate DACH breakdown.

5 OpenAI, "ChatGPT Search", accessed 28 July 2026, OpenAI Help Center. Also OpenAI, "Introducing ChatGPT search", 31 October 2024, OpenAI. Sources describe automatic and manual search triggering, query rewriting, source display and possible location and memory context on the named consumer surface.

6 Google Search Central, "AI features and your website", and Google Search Help, "AI Mode in Google Search", accessed 28 July 2026, Google Search Central and Google Search Help. Sources describe the distinction between AI Overviews and AI Mode, conditional serving, query fan-out and eligibility. No full ranking formula.

7 Microsoft, "How web search works in Microsoft 365 Copilot Chat and agents", accessed 28 July 2026, Microsoft Support. Microsoft Bing, "Introducing Copilot Search in Bing", April 2025, Bing Blog. Sources describe distinct Microsoft surfaces and policy-dependent web search.

8 Perplexity, "What is Pro Search?", updated 21 July 2026, Perplexity Help Center. Perplexity, "Sonar API Quickstart" and "Pro Search Quickstart", accessed 28 July 2026, Sonar API and Pro Search. Sources describe differences between modes and API products.

9 Anthropic, "Enable and use web search", and "Web search API", 7 May 2025, accessed 28 July 2026, Claude Help Center and Anthropic. Sources describe web search as an activatable, partially policy-controlled tool state.

10 Google, "View related sources from Gemini apps", accessed 28 July 2026, Gemini Apps Help. Google AI for Developers, "Grounding with Google Search", accessed 28 July 2026, Google AI for Developers. Sources describe the display of sources in the app and structured grounding metadata on the named API surface.

11 OpenAI, "Publishers and Developers FAQ", accessed 28 July 2026, OpenAI Help Center. Documents describe OAI-SearchBot as a technical precondition for possible consideration in ChatGPT Search. No guarantee of crawl, indexing, retrieval, citation or recommendation.

12 Nelson F. Liu, Tianyi Zhang and Percy Liang, "Evaluating Verifiability in Generative Search Engines", arXiv 2304.09848, 2023, arXiv. Michael Wornow et al., "SourceCheckup", Nature Communications 16, 2025, Nature Communications. Used only for the methodological distinction between citation presence, assignment and actual support.

13 Google Search Help, "Understand and manage your location when you search on Google", Google Search Help. OpenAI, "ChatGPT Search", OpenAI Help Center. Anthropic, "Does Claude use my location?", Anthropic Privacy Center, all accessed 28 July 2026. Documents describe possible, surface-dependent location signals.

14 Google Business Profile Help, "How Google sources and uses information in Business Profiles", Google Business Profile. Google Maps Platform, "Places API Overview", Google for Developers. Google Business Profile Help, "Create a bulk upload spreadsheet for Business Profiles", Google Business Profile. Google Business Profile APIs, "Manage locations at scale", Google for Developers, all accessed 28 July 2026.

15 National Institute of Standards and Technology, "AI RMF Core, Measure", esp. Measure 2.1 to 2.5, accessed 28 July 2026, NIST AI RMF. American Association for Public Opinion Research, "Disclosure Standards", accessed 28 July 2026, AAPOR Disclosure Standards.

16 Gili Lior et al., "ReliableEval. A Recipe for Stochastic LLM Evaluation via Method of Moments", Findings of EMNLP 2025, ACL Anthology. Settings-specific parameters not transferred as a universal prompt or run count.

17 Andong Hua et al., "Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs", EMNLP 2025, ACL Anthology. Counter-finding on evaluation-induced prompt sensitivity. Does not show that output variation is generally an artefact.

18 Aaron Chatterji et al., "How People Use ChatGPT", NBER Working Paper 34255, September 2025, NBER. Several authors were employed by the ChatGPT operator at time of publication. First-party data used only to describe broad usage categories.

19 The two anonymised commercial illustrations come from the documented claim sample of the essay. Originals, wording and access date are internally archived. Vendor, product, metric and link details removed to prevent reidentification.

20 OECD, European Union and European Commission Joint Research Centre, "Handbook on Constructing Composite Indicators", 22 August 2008, esp. chapters 5 and 6, OECD. General rules on selection, normalisation, weighting, aggregation and sensitivity analysis.

21 Google Search Central, "Introducing Search Generative AI performance reports in Search Console", 3 June 2026, accessed 28 July 2026, Google Search Central. Documented URL impressions apply to the named Google features. Not a full cross-surface exposure measurement.

22 Google Analytics Help, "About traffic-source dimensions", accessed 28 July 2026, Google Analytics. Referral and source dimensions carry engagement measurement when a source signal is present.

23 European Commission Joint Research Centre, "Counterfactual impact evaluation", accessed 28 July 2026, JRC. Institutional methodology treats a counterfactual as a precondition of any isolated impact claim.

24 Own source-review synthesis for Essay 05, research cut-off 28 July 2026. Reviewed the scientific, institutional, technical and commercial sources documented in the research corpus on synthetic output measurement, first-party impressions, referral analytics and outcomes. Finding is time-bound and not a systematic meta-analysis.

65 The 30 / 90 / 180 day structure is a disclosed editorial planning convention from the research synthesis of Essay 05. It is neither an empirical industry standard nor a minimum duration, budget assumption or promised effect latency. Real gates must be derived from data state, portfolio, measurement design, organisation and decision.

66 The TCO schema is an editorial synthesis of the work and cost categories documented in the research. Public vendor, product, tariff, price and billing details remain in the private evidence edition and are not disclosed here. Schema is not a market, personnel or budget benchmark.

Sources

References

References

Whitepaper PDF https://www.localsearch.ch/app/uploads/2025/08/Whitepaper_KMU-Digital-Pulse-2025_DE.pdf

Reuters Institute https://reutersinstitute.politics.ox.ac.uk/generative-ai-and-news-report-2025-how-people-think-about-ais-role-journalism-and-society

Pew Research Center https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/

SSRN https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6513059

arXiv https://arxiv.org/abs/2602.18455

INFORMS https://pubsonline.informs.org/doi/10.1287/mksc.2025.0489

OpenAI Help Center https://help.openai.com/en/articles/9237897-chatgpt-search

OpenAI https://openai.com/index/introducing-chatgpt-search/

Google Search Central https://developers.google.com/search/docs/appearance/ai-features

Google Search Help https://support.google.com/websearch/answer/16011537?hl=de-DE

Microsoft Support https://support.microsoft.com/en-us/Microsoft-365-Copilot/how-web-search-works-in-microsoft-365-copilot-chat-and-agents

Bing Blog https://blogs.bing.com/search/April-2025/Introducing-Copilot-Search-in-Bing

Perplexity Help Center https://www.perplexity.ai/help-center/en/articles/10352903-what-is-pro-search

Sonar API https://docs.perplexity.ai/docs/sonar/quickstart

Pro Search https://docs.perplexity.ai/docs/sonar/pro-search/quickstart

Claude Help Center https://support.claude.com/en/articles/10684626-enable-and-use-web-search

Anthropic https://www.anthropic.com/news/web-search-api

Gemini Apps Help https://support.google.com/gemini/answer/14143489?co=GENIE.Platform%3DDesktop&hl=de

Google AI for Developers https://ai.google.dev/gemini-api/docs/google-search

OpenAI Help Center https://help.openai.com/en/articles/12627856-publishers-and-developers-faq

arXiv https://arxiv.org/abs/2304.09848

Nature Communications https://www.nature.com/articles/s41467-025-58551-6

Google Search Help https://support.google.com/websearch/answer/179386?hl=en

Anthropic Privacy Center https://privacy.claude.com/en/articles/11186740-does-claude-use-my-location

Google Business Profile https://support.google.com/business/answer/2721884?hl=en

Google for Developers https://developers.google.com/maps/documentation/places/web-service/op-overview

Google Business Profile https://support.google.com/business/answer/3370250?hl=en

Google for Developers https://developers.google.com/my-business/content/manage-locations

NIST AI RMF https://airc.nist.gov/airmf-resources/airmf/5-sec-core/

AAPOR Disclosure Standards https://aapor.org/standards-and-ethics/disclosure-standards/

ACL Anthology https://aclanthology.org/2025.findings-emnlp.594/

ACL Anthology https://aclanthology.org/2025.emnlp-main.1006/

NBER https://www.nber.org/papers/w34255

OECD https://doi.org/10.1787/9789264043466-en

Google Search Central https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports

Google Analytics https://support.google.com/analytics/answer/15612152?hl=en

JRC https://joint-research-centre.ec.europa.eu/projects-and-activities/counterfactual-impact-evaluation_en

Correspondence

If this piece reflects a question you are weighing, a short note is a good place to begin.