← The libraryText edition
Format
Essay
Reading time
11 min
Reading level
Considered
Published
13 September 2026
Topics
Executive Decision Making
Technology Strategy
Artificial Intelligence
Leadership

Executive Essay 03

The Architecture of Better Questions

On the assumptions a technology decision asks us to accept

An evaluation can identify the best answer to a brief while leaving its diagnosis unexamined. This essay considers when leaders have reason to revisit that diagnosis, how a change of purpose differs from a correction, and what can justify proceeding with uncertainty still unresolved.

Part of Executive Essays, standalone arguments developed from recurring patterns in organisational practice.

Imagine a proposal awaiting your approval. The team has compared AI applications for customer enquiries, tested their responses and worked through the cost of connecting them to existing systems. The recommendation is clear. You can follow the reasoning from the evaluation criteria to the preferred option, including the compromises it requires.

There is enough here to have a serious discussion about the purchase. Whether there is enough to approve it depends partly on something that may sit several pages behind the comparison, or in an earlier document altogether. What led the business to treat these enquiries as a problem for an AI application to solve?

The team may already have established why an AI application is a reasonable response, having examined when customers lack advice and ruled out simpler changes for good reasons.

But the comparison itself cannot tell you that. It can establish which application best meets the brief while leaving the diagnosis inside that brief untouched. You are being asked to commit resources to an answer. The proposal should also explain why the business has chosen to address this problem in this way.

Before the comparison

A brief asking which system should handle customer enquiries already gives part of the work its direction. It treats handling the enquiries as the relevant intervention. If some arise because customers cannot understand a policy, changing that policy would sit outside the comparison unless the brief admitted it. This example shows how an exclusion can enter the work before anyone scores a supplier.

Problem definition is more substantial than wording. The same projected expenditure can be presented as an amount spent or as a saving against the current budget. That changes how the option is described without changing what is being proposed. Deciding that a service problem requires a new system determines what will be investigated. A justification assembled after choosing the system does something else again. It may make the choice intelligible without showing that the underlying assumption was examined when it could still have changed the decision.

In their theoretical account of strategic problem formulation, Markus Baer, Kurt Dirks and Jackson Nickerson connect restricted problem formulation with limits on the search for solutions. Their work offers a conceptual basis for examining those boundaries, rather than evidence that a particular questioning procedure improves project outcomes. 1

The concern follows from the Foundation Series’ examination of judgement, organisational memory and time. Here, the question is what those capacities are being asked to examine in the first place.

I would place responsibility for that choice with the leadership approving the commitment. This does not require executives to redo specialist analysis. It requires them to understand why the investigation has these bounds, particularly where an assumption could exclude a materially different response. A narrow brief can be the mark of work already done well. Its supporting reasons should survive the shortening of the document.

Nor does a sound definition settle technical suitability or the economics of delivery. Integration may remain the hardest problem. The more limited concern is that, where the diagnosis is untested, rigorous analysis of the permitted options can leave the original difficulty unresolved. Revisiting the diagnosis offers no assurance of a better outcome. Evidence of an application’s performance must therefore be read with attention to the decision it actually helps to answer.

Improvement against what?

For an executive considering AI, a positive result elsewhere can make a proposal more credible. Its relevance depends on what changed for the people in that study and what they would otherwise have received. The comparison is part of the substance of the finding.

In a preprint on online retail, Lu Fang and colleagues report experiments on an unnamed platform. Their sales chatbot produced a statistically significant revenue advantage over a condition without live advice, where customers received an automatic message that service was unavailable. Revenue here meant product spending per consumer. An additional comparison with human replies found no statistically significant revenue difference. That result establishes neither equivalence nor non-inferiority, and says nothing by itself about equal quality of advice. 2

The evidence comes from short experiments on one platform. During the project, two authors served as consultants to the company and a third was employed by it. These qualifications belong alongside the finding when considering how far it travels. 2

For our hypothetical approval, the useful distinction is between making advice available where it was absent and changing who provides advice that already exists. An organisation facing an availability problem might reasonably find the first comparison relevant. A proposal to replace an established service needs evidence against that service. Carrying the revenue advantage from one comparison into the other would change the claim supporting the investment.

A brief that attributes the difficulty to staff capacity directs attention towards tools that can handle more enquiries. If the difficulty instead lies in inconsistent guidance, the investigation needs to include how that guidance is produced and maintained.

The study provides no account of how management arrived at the project brief. It gives us a result whose meaning depends on its baseline. Even a promising finding leaves the adopting business to establish whether the same gap exists among its own customers and whether filling it would justify the proposed commitment.

A separate study by Erik Brynjolfsson, Danielle Li and Lindsey Raymond helps locate a different boundary. Their field study of a staggered introduction of AI assistance in a software company’s customer service operation estimated an increase in resolved enquiries per hour, with effects differing across workers. Employees remained responsible for replies. They could edit suggestions, use parts of them or ignore them. The main evaluation was not a randomised trial. 3

Here the limit concerns the task being transferred. A finding obtained while people continue to handle enquiries does not establish what happens when that responsibility is removed. It supports consideration of the arrangement studied. It cannot, on its own, answer a staffing question about complete replacement.

The two studies leave us with distinct obligations when using their findings. We need a relevant comparison and an accurate account of the work that remains with people. Neither study tests whether reframing produces better projects. A broadly encouraging AI result cannot substitute for a precise account of the proposed change. Once that change is clear, there is a further choice about what the system will treat as success.

What the target represents

That choice has consequences beyond the way a project is reported. A system can be trained to predict a measurable quantity which only imperfectly represents the purpose for which its predictions will be used.

Ziad Obermeyer and colleagues examined a predictive algorithm used in the US health system to help select patients for additional care. It predicted total medical spending, which served as a proxy for health need. In the population studied, Black patients were sicker than White patients at the same risk score, as assessed through chronic conditions and clinical measures. At comparable measured health, spending on Black patients was lower. To explain the spending differences, the authors drew on research concerning barriers to access and racially unequal care. Their analysis did not separately establish the causal contribution of each of those channels. 4

The distinction between expenditure and need matters here because it affects which patients are identified for help. Reading the case simply as a lesson about choosing a better metric would strip away the racial inequality in the US health system that gives the mismatch its substance. Expenditure records what was spent within that system. Using those spending patterns to identify need can carry existing disparities in care into decisions about who receives additional support.

The translation of a purpose into a predictive target deserves examination in its own right. It is related to problem definition but does not stand for the whole of it. A company may have a defensible purpose and still choose a poor measure of progress towards it. Changing the measure might preserve the purpose. A broader reformulation could instead introduce a different purpose, which requires a different justification.

A different diagnosis, or a different purpose?

Return to the hypothetical customer service proposal. Suppose the business wants customers to receive dependable answers sooner, and examination now supports the explanation that employees spend time resolving inconsistent guidance before they can answer. Correcting the guidance becomes a relevant option. The diagnosis and proposed intervention have changed, while the desired outcome remains dependable answers delivered sooner.

Now suppose someone proposes reducing the cost of customer contact by moving more customers to self-service, while accepting that some will find it harder to obtain an answer. There may be a commercial case for doing so. It changes what the organisation is willing to provide, however, and whose inconvenience it is prepared to accept. Calling it a clearer definition of the same problem would conceal a choice about the service itself.

This is where enthusiasm for reframing needs restraint. A new question can offer an attractive objective without explaining the condition that prompted the original proposal. It may shift expenditure to another department or effort to the customer. The person presenting it should be able to say whether the diagnosis has changed or whether a different outcome is now being sought. The authority to make the latter choice belongs with those accountable for its consequences.

Measures need similar care. If elapsed time becomes the sole measure of dependable service, improving the score may leave the reliability of the answer unexamined. Revising that measure need not alter the service promise. Deliberately reducing the promise does. The distinction gives the discussion somewhere concrete to go when two participants appear to disagree about the same brief but are actually defending different outcomes.

It also protects decisions already made. A contractual obligation or an explicitly agreed service commitment cannot be dismissed as a failure of imagination. Those constraints may be open to renegotiation, but that is a further decision with its own consequences. Reopening a brief should make such choices visible. It should not quietly grant the person proposing a new formulation permission to make them.

Enough to proceed

There is a point at which another discussion of the question becomes an obstacle to answering it. People need a sufficiently settled assignment to build and operate something. An executive who repeatedly reopens the purpose can impose uncertainty on everyone else while describing it as intellectual care.

I would reopen the brief where a consequential assumption lacks support or conflicts with what the organisation is seeing, and examining it could change the decision. If customers are waiting for policy clarification, a proposal based on insufficient response capacity deserves examination because a different diagnosis could change the investment. The mere possibility of describing customer service more broadly does not justify another investigation. A colleague asking to reopen the brief owes the team an account of what might change as a result.

The grounds for closure are a matter of judgement. They are not supplied by the studies discussed here. The responsible executive needs to consider what can still be learned before the commitment, at what cost, and with what bearing on the choice. The significance of an error matters. So does the disruption caused by keeping a team in preparation while the existing problem continues.

This examination needs someone who can bring it to an end. That person need not possess every answer, but must be able to establish which doubt is being investigated and decide whether the resulting evidence is sufficient. If records already show why an alternative was ruled out, recovering that reasoning may settle the matter. Asking the team to repeat it from the beginning would add work without necessarily adding understanding.

Where the evidence remains incomplete, a bounded trial may be useful. It must expose the uncertainty that matters. Demonstrating that an assistant can generate a fluent reply would not test whether inconsistent guidance is the source of delay. A trial also deserves scrutiny if the required integration or contractual commitment would effectively make the decision it is supposed to inform. Calling an activity a pilot does not make its consequences reversible.

Further inquiry can reasonably end when the grounds are sufficient to confirm or change the definition. It can also end when no proportionate investigation before the decision is likely to resolve the remaining uncertainty. In that situation, proceeding means accepting the possible consequences of being wrong. A deadline explains why a decision is needed now. It supplies no additional evidence that the diagnosis is correct. Depending on the exposure, declining the proposal may be more defensible than either approving it or commissioning another study.

Once the problem is adequately defined, the work may properly return to a technical constraint. Unreliable data still need attention. A difficult integration needs engineering capacity, and an application the business cannot afford to run does not become viable through a better question. Skills and implementation can be the principal obstacles. Further examination of the brief needs a reason when the available evidence points to a delivery problem.

Closing the inquiry should nevertheless leave a usable account of what remains uncertain. If later evidence contradicts a premise on which the commitment depended, there is a reason to revisit it. The existence of sunk effort should not settle that judgement, although contracts and technical dependencies still affect what can be changed. Equally, a disappointing outcome alone does not establish that the earlier definition was unreasonable. The decision deserves examination against what could have been known then, rather than a reconstructed brief that makes the outcome look inevitable.

The approval

The proposal returns to the table. It may recommend exactly the same application. Nothing in this argument requires a new supplier, a wider scope or an additional meeting to prove that leadership has exercised judgement. The earlier narrowing may have been entirely sound, with reasons that were simply absent from the document you first read.

If the examination did change the proposal, the change should be intelligible to the people expected to carry it out. They need to know whether they are addressing the same difficulty through a different intervention or have been given a different purpose. Otherwise, an apparent clarification at the approval table can leave contradictory expectations in the work that follows.

Your approval cannot certify the future result. It can express a considered willingness to act on this account of the problem, with these limits to what is known. That is a more demanding commitment than accepting the recommendation with the highest score. It also permits the team to proceed without pretending that every uncertainty has been removed.

Before the signature, there is one question I would want the proposal to make answerable. Which assumption about the problem are we accepting here, and what gives us sufficient reason to act on it?

Sources

References

1. Baer, Markus, Kurt T. Dirks and Jackson A. Nickerson (2013, first published online 2012). “Microfoundations of strategic problem formulation.” Strategic Management Journal, 34(2), 197–214. DOI 10.1002/smj.2004. University record and abstract.

2. Fang, Lu, Zhe Yuan, Kaifu Zhang, Dante Donati and Miklos Sarvary (2026). Generative AI and Sales Productivity: Field Experiments in Online Retail. Preprint, arXiv:2510.12049v6. Version dated 29 June 2026, manuscript dated 30 June 2026. Version record. Full text v6. See §3.1, Table 4 and Appendix C1, Table C1.

3. Brynjolfsson, Erik, Danielle Li and Lindsey Raymond (2024). Generative AI at Work. arXiv:2304.11771v2. Version submitted 6 November 2024, manuscript dated 7 November 2024. Version record. Full text v2. See §§2.3, 3.1–3.3, 4.1–4.2 and 5.1.

4. Obermeyer, Ziad, Brian Powers, Christine Vogeli and Sendhil Mullainathan (2019). “Dissecting racial bias in an algorithm used to manage the health of populations.” Science, 366(6464), 447–453, 25 October 2019. DOI 10.1126/science.aax2342. Published article hosted by the FTC. See “Mechanism of bias” and “Problem formulation”, pp. 449–451.

Correspondence

If this piece reflects a question you are weighing, a short note is a good place to begin.