- Format
- Essay
- Reading time
- 15 min
- Reading level
- Considered
- Published
- 14 September 2026
- Topics
- Executive Decision MakingDigital CommerceArtificial IntelligenceTechnology Strategy
Executive Essay 04
When AI Visibility Becomes a Budget Decision
On the evidence needed to commit money and people
A useful visibility report can support an investigation without settling the case for expenditure. This essay considers the evidence a retailer needs to correct product information, test an improvement or make a continuing commitment.
Part of Executive Essays, standalone arguments developed from recurring patterns in organisational practice.
A proposal to improve AI visibility can contain several quite different requests. Some product information needs correcting. A set of buying guides might be worth commissioning. There is also a suggestion that the whole catalogue should be rewritten, supported by a new monitoring contract.
The report accompanying the proposal may be useful. It cannot, by itself, establish that all three requests deserve funding.
Consider a figure that would sit comfortably near the front of a budget presentation. Google documents a circumstance in which Merchant Center reports 100 per cent share of voice because an account has no defined competitors available. Treating that figure as evidence of market dominance would give the report rather more credit than its definition warrants. Google Merchant Center documentation
There need be nothing wrong with the calculation for the spending decision built upon it to be poorly founded. More detailed information about products, search intent and competitors makes the investigation more useful. It does not remove the need to explain why a proposed change should work, or why its expected benefit warrants the cost.
My earlier essay, AI Visibility. What Can Be Measured and What Is Merely Sold, examined the quality of these observations and their limits as a basis for management. A retrieval, a mention, a citation, a visit and a commercial outcome are different events. A monitoring system does not automatically establish the causal connections between them.
Here I want to examine the commitment that follows. Generative engine optimisation, or GEO, covers work intended to improve visibility in generative search and answer systems. Within that description sit decisions with very different consequences. Correcting a product specification and establishing a permanent team should not require the same case.
Demanding a complete account of commercial impact before making any useful correction would bring ordinary work to a halt. Accepting every improvement in a visibility score as a reason to spend more would be expensive in a different way. The evidence needs to be adequate for the decision being made.
A clearer view of a limited field
Google's AI performance report provides information on search intent, terms and product attributes for organic shopping visibility in AI Mode and AI Overviews. Its share of voice relates a brand's AI impressions to the combined impressions of that brand and the competitors included in the report. Availability is limited to English-language queries for Merchant Center accounts in Australia, Canada, India, New Zealand and the United States. It is therefore not generally available to retailers in Germany, Austria or Switzerland. Google's report scope and availability
That additional context can turn a general concern about being overlooked into a more specific investigation. A retailer may be able to identify particular products or information needs where its visibility is weak. This is a worthwhile advance.
The relevance of the finding still depends on what has been observed. The queries, countries, interfaces and competitors represented in a report need to correspond sufficiently to the business considering the expenditure. A carefully defined US metric may tell a retailer trading only in Germany very little about its immediate opportunity.
Direct platform data has the advantage of showing part of what happens within the platform itself. Extending that observation to another market, another interface or a business outcome requires a further argument.
Better data can therefore narrow the search for a problem. It can point towards work worth investigating. The choice of remedy and the case for paying for it remain to be made.
Three proposals for one catalogue
Imagine a specialist running-shoe retailer. The example is hypothetical. The business regularly tests a fixed selection of German-language questions in AI answer systems and finds that competitors appear more frequently in responses about shoes for wide feet. Its selection of questions, or prompt panel, covers chosen enquiries. It does not represent German purchasing behaviour as a whole.
An examination of the catalogue reveals several different issues.
For some shoes, the manufacturer's confirmed width information is available internally but missing from the shop. The immediate task is to add it accurately and check that the product page, feed and filters agree. Google's guidance for its AI search features continues to emphasise the established fundamentals of accessible, consistent information. It does not require additional AI files or a special schema for a page to be eligible as a supporting link in AI Overviews or AI Mode. Google Search Central
Elsewhere in the catalogue, the specifications are complete. What remains unclear is whether customers understand the differences between the shape of the last, the stated width and the resulting fit. A comparison guide could help. The questions customers actually ask and the team's knowledge of the products should determine what it explains. It must not promise a fit the retailer cannot substantiate.
Now suppose a supplier proposes rewriting the entire catalogue using a new GEO method. That is a much larger undertaking than filling the identified gaps. The retailer would need to understand why this particular method should produce an additional benefit and why the work deserves priority over other improvements.
All three proposals might be described as product data optimisation. They should still be considered separately. The missing information can be corrected, a limited set of comparisons can be tested, and the broader production proposal can be assessed on its own merits. The urgency of fixing a known omission does not establish the case for rewriting everything around it.
What each proposal needs to justify
The example suggests a practical distinction between correcting an error, testing an improvement and making a continuing commitment. I offer it as a way of organising the decision, not as an empirically validated industry standard.
The width specification needs correcting because customers should be given reliable information. Requiring proof of a benefit to AI visibility would add a condition unrelated to that immediate responsibility.
The proposed fit guidance has a less certain case. Customer questions and a pattern of competitor mentions provide reasons to investigate. They do not establish that the guidance will address the cause of the visibility gap. A trial can be justified without treating its hoped-for result as a foregone conclusion.
The catalogue-wide commission is different again. Like a permanent team or a longer software contract, it competes with other uses of the same resources. A successful example would not normally be enough. The business needs a reasonable expectation of repeatable benefit, a credible account of costs and a basis for reconsidering the commitment.
The more costly an error would be, and the harder it would be to reverse, the stronger that account should become. The calculation must also include the possible cost of waiting. If relevant demand is moving towards a new route to purchase, delay may have a price. Its likely importance to this particular business still needs explaining. The pace of the market is context, not a completed investment case.
This allows a retailer to act without first solving the behaviour of a language model. What matters is recognising the scale of the decision and the evidence it calls for.
What the recommendation has established
Research can help make the transition from observation to recommendation more explicit. In a paper submitted on 10 September 2026, Spandan Ghose Chowdhury presents Agentic Share-of-Search, a system that analyses product mentions and recommends catalogue improvements. Across 100 synthetic tests, it identified the deliberately altered signal in 39 cases. Original paper
The design of those tests determines what the result means. Attribute signals were changed while visibility outcomes were held constant. The task was to recover the altered signal. It could not establish whether acting on the recommendation would increase visibility. The author explicitly describes the diagnosis as association-based and distinguishes it from an estimate of intervention impact. Full paper, diagnostic method and evaluation
This makes the route to a recommendation open to examination. Further work is needed before it can support an investment decision.
Our hypothetical retailer would first have to establish whether the diagnosed problem exists in its own catalogue. It would then need a plausible account of how the proposed change could address it. Implementation and a suitable evaluation would be required to determine whether an outcome improved beyond what would otherwise have been expected.
An advisory proposal should be clear about which of these steps its evidence supports. A convincing diagnosis has value. The effectiveness of the action proposed in response needs an assessment of its own.
Give the experiment a decision to inform
For the running-shoe retailer, a limited trial could establish whether properly researched fit comparisons warrant a larger editorial commitment. Before publishing them, the business should decide what the trial is intended to resolve.
Helping a customer judge fit, earning more mentions in an answer system and generating additional orders are different research questions. They may be related, but a favourable result on one does not answer the others.
The comparison also needs care. A change between the weeks before and after publication could reflect seasonal demand, a competitor's offer, an alteration to the model or a campaign running at the same time. Counterfactual evaluation asks what would have happened without the intervention. The European Commission's Joint Research Centre explains this principle in its guidance on impact evaluation. Applying it to the retailer's trial here is my methodological interpretation. Joint Research Centre
Known errors in the specifications would be corrected before the trial. For the additional comparisons, the retailer could identify suitable product or category groups and, where appropriate, randomly assign them to earlier or later introduction. A visitor test within the shop addresses the effect of a presentation on those visitors. It does not demonstrate that an external answer system recommends the content more often.
The groups need to suit the question. Two categories of similar size may differ substantially in seasonality, margin and demand. A category description can affect several products at once, while links can carry an effect beyond the pages directly changed. Where those connections matter, treating every product page as an independent test unit would be misleading.
Repeated observation is also necessary when examining AI answers. The study Don't Measure Once describes why visibility should be assessed as a distribution across queries and time rather than inferred from a single result. Schulte, Bleeker and Kaufmann Repeating a question a hundred times does not, however, produce a sample of a hundred independent customers. It helps describe variation within the chosen test. It does not measure the size of the market asking that question.
There is a related discipline in choosing the questions themselves. The retailer should not narrow the panel after the event to the wording on which it performs well. Some questions could be kept out of the content development process and used later for evaluation. If improvement appears only on the formulations used to shape the copy, its wider relevance remains uncertain. Holding questions back does not make the panel representative. It does make it harder for the evaluation simply to reward the text for matching its own brief.
The appropriate duration depends on the number of relevant events, delays in updating and the size of the effect the test could detect. If an observed system changes its model or the presentation of its answers substantially during the trial, comparability needs to be reassessed.
A retailer with very few relevant purchases may be unable to establish a reliable revenue effect through a short trial, however carefully designed. That limits the conclusion available. It does not automatically rule out the work.
Useful content and the question of credit
Suppose the fit comparisons prove helpful in our hypothetical trial. Customers find a suitable variant more quickly, and the retailer sees indications of fewer avoidable enquiries. Over the same period, its products also appear more often in some AI answers.
The content could be helping customers regardless of how they reached the shop. The extra mentions could result from the same change, from other influences or from ordinary variation. Their appearance in the same period would not establish a causal connection.
This distinction has consequences for expenditure. A sufficiently well-supported benefit within the shop could justify the fit guidance. It would not, on its own, support the claim that a particular GEO method had generated additional revenue. The editorial work might deserve continuation while the case for a separate GEO contract remained unresolved.
Where a change is intended to improve product advice, conventional search and AI discovery together, allocating the benefit cleanly between channels may be difficult. I would first assess whether the work as a whole is worth doing. I would credit a particular channel only to the extent the evaluation supports it.
That leaves room to retain useful work without attaching an unsupported explanation to its success.
Following a sale does not establish an additional sale
The methodological note accompanying my earlier essay considered the missing links between retrievals, answers and subsequent visits. Suppose those links became clearer within a particular system. A retailer might be able to connect a product recommendation, the resulting visit and a purchase within the same customer journey.
That would be a considerable improvement in what could be observed. It would still leave open whether the customer would have purchased without the recommendation. A connected sequence of events establishes an observed route. The additional contribution of the intervention is a further question.
Commercial value also depends on what follows the order. More attributed purchases could come with more returns, heavier discounting or higher service costs. Fewer, better-informed buyers might make a more valuable contribution. These are possible outcomes in the hypothetical example, not empirical findings being claimed here.
For the shoe retailer, contribution margin, returns and avoidable service costs would therefore belong in the assessment. A visibility measure might help explain a change. Whether that change benefits the business depends on its commercial consequences.
Investing before every effect can be isolated
A smaller retailer may be unable to isolate the additional effect of GEO at a reasonable cost for some time. It still has to decide whether an experiment is worth funding.
A limited commitment can be defensible under uncertainty. There should be a plausible account of how the work could help, some concrete benefit if only part of the expectation is realised, and a possible loss the business can afford. Our retailer could commission a small number of comparisons answering frequent questions about fit. Their use within the shop would already have a plausible purpose. Their subsequent appearance in AI answers would be an additional expectation to investigate.
A simple calculation can make the scale of that expectation visible. Assume, solely for this example, total costs of €6,000 over a defined period. At an average incremental contribution of €30 per additional order, the work would need 200 additional orders within that period to cover its costs. The €30 would already have to account for returns and relevant variable costs. The €6,000 would need to include all costs of the work during the period, including maintenance and evaluation.
Under those assumptions, the calculation establishes how many additional orders would be required. It does not predict that they will arrive. If relevant demand is small, the scale alone may count against the proposal. Where demand is plausible, it makes the expectation more concrete and helps frame an appropriately sized trial.
A larger commitment should not depend entirely on the most favourable assumptions. The retailer should consider whether the expenditure remains defensible if the effect is smaller, the content costs more to maintain or AI discovery brings fewer visitors than hoped. Scenarios do not replace evidence. They expose how dependent the decision is on expectations that have yet to be established.
A business may also choose to move early because it anticipates a shift in the market. It should describe that choice as a reasoned risk and limit its exposure accordingly. A more detailed report does not turn a bet into a demonstrated return.
What further measurement would change
Measurement itself consumes resources. Checking data, interpreting findings and discussing them take time. For a smaller organisation, that effort needs to be proportionate to the expected value of making a better decision.
The practical question is what the retailer would do differently with a more precise result. If greater precision would change neither the scope of the work nor the decision to continue it, the value of obtaining it is limited.
The response to possible findings should be considered before the trial begins. A convincing result may support expansion. Contradictory findings first need an explanation. An inconclusive result warrants more observation only if that observation is likely to help resolve the decision. Leaving an unsuitable trial running for longer does not repair its design.
Failure to demonstrate an effect is not automatically proof that there is none. It is also insufficient grounds for an indefinite commitment. In the shoe example, accurate specifications and useful fit guidance could remain in place while a costly expansion or an uninformative monitoring service was stopped. The assessment of the content need not settle the value of every service surrounding it.
For the person responsible, four questions should remain answerable. Which decision is this observation meant to change? What additional effect do we expect from the proposed work? What evidence is sufficient given its cost and reversibility? And what result would lead us to change or stop it?
Returning to the budget discussion, the retailer would then have something more useful than an explanation of a percentage. It could say which information will be corrected, which extension merits a trial and why a larger commission will have to wait. It could also explain what would change that judgement.
That is the practical value I would ask of better GEO data. It should help a business make a more considered commitment of money and people. The size of the commitment remains a matter of commercial judgement.
Editorial note. Prepared on 14 September 2026. The retailer example, including all assumed outcomes and figures, is hypothetical. The proposed decision framework and the application of evaluation methods to GEO are the author's interpretations. This essay continues the earlier work on AI visibility. It does not report a study using original retailer data.