Product management with the means to build

How AI expands the questions a product manager should be able to answer

In this essay

Early in Arrival, a military officer asks Louise Banks to interpret a recording of an unfamiliar language. He wants an answer from the material he has brought. She needs a way to investigate what the material means. In the screenplay, she explains the condition under which she could do the work. “To do this right, I need to be there. Interacting with them.”1

The difficulty is familiar to anyone who has received a confident request for a product whose underlying problem remains unclear. The requested output may be straightforward to describe. Establishing whether it would help requires access to the situation that produced the request, a way to distinguish plausible explanations, and sometimes an intervention that makes a previously invisible constraint observable.

Good product management has always involved that work. It would be a mistake to describe the traditional profession as writing documents and coordinating other people, then claim that AI has invented the need for customer understanding or technical curiosity. The more consequential change is that a product manager can increasingly carry an investigation into forms that once required a substantial commitment from other specialists. A working interaction, a data analysis, or a small executable system can become part of discovering what the product should be.

That changes the standard we can reasonably expect. When the cost of testing an interpretation falls, leaving it as an ambiguous sentence becomes harder to justify. When software can act on a user’s behalf, the meaning of a successful action becomes part of the product definition. And when models can generate convincing arguments for almost any proposed direction, a product manager needs a better way to decide which argument deserves investment.

The opportunity is to extend product reasoning farther into the work: from understanding a need to constructing an experience, examining its behavior, and establishing whether it changes the outcome that justified building it. This could make the role more technical, but its larger consequence is a different relationship between what a product manager believes and what they can do to find out.

1. A request is the beginning of the problem

Consider a hypothetical company making software for independent appliance-repair businesses. Its customers ask for an AI booking assistant because employees spend time answering calls, collecting details, and finding appointments. It would be easy to begin with a conversational interface, a calendar connection, and a target for the number of bookings completed without assistance.

A more useful investigation would follow what happens after a booking. Does the technician arrive with the right part? Was enough information collected to identify the appliance? Could the customer authorize the likely cost? Was the appointment actually available, or did another channel allocate the same time? The apparent administrative burden may be one symptom of a less visible coordination problem.

Different customers could need different products. A business handling predictable maintenance has a different information problem from one receiving urgent calls about poorly documented faults. A shop whose owner answers every call may value relief from interruption; a larger operator may be losing money on visits that fail because the necessary equipment was not available. Segmenting by those circumstances helps identify the intervention that matters. Calling both groups “small businesses” does not.

AI can make this investigation broader and more precise. Given authorized records, it can help trace failed visits back through intake messages, bookings, and inventory events, identify competing explanations, and propose questions for conversations with the people doing the work. It can also help inspect the cases that do not fit the emerging account. The important distinction is between evidence about these customers and a plausible simulation of how a customer might respond. A generated persona can expose an assumption; its agreement cannot establish demand.

The discipline remains first-principles thinking, but the means of applying it expand. We ask which observations support the explanation, which steps are inferences, and what would look different if another explanation were true. A business reporting “too many calls” might be experiencing unclear pricing, failed appointments, or a genuinely cumbersome intake process. Automating the call addresses those problems differently. The earliest useful test may concern the reason people call again.

There is also a strategic question. Why is this company well placed to solve the problem? A generic assistant may produce a pleasant conversation, while the existing software provider can connect the conversation to parts availability, technician skills, and actual customer commitments. That advantage has value only if those records are sufficiently reliable and accessible. Possessing a database is not the same as being able to make a dependable promise from it.

An AI-era product manager should be able to pursue this chain far enough to change the investment. That could mean demonstrating that a booking assistant is exactly the right first product. It could mean discovering that a small change to the information collected would do more. It could also reveal a much larger opportunity to coordinate the repair itself.

2. Vision has to survive becoming an experience

A diagnosis explains what is wrong. Product vision explains a better way of living or working that becomes possible when the problem is solved. AI expands the range of experiences we can attempt, but it also makes it easy to substitute a description of intelligence for a description of a product.

For the repair business, “an intelligent assistant that understands every customer” tells us little about what changes. Imagine instead a customer photographing an appliance label and describing a fault. The service identifies information it still needs, checks whether the business can support the repair, and offers a visit window that accounts for both technician availability and the expected part. The customer sees the price of diagnosis and the conditions under which further work would need approval. If the part is uncertain, the service says so before offering an appointment whose usefulness depends on it.

Now imagine what the business receives. The technician can see what was reported, what remains uncertain, and why the appointment was scheduled. A parts delay changes the promise before someone leaves for an avoidable visit. A customer’s authorization is attached to the work it permits. The useful experience continues after the conversation ends.

That vision creates several possible product directions. One could improve intake while leaving scheduling decisions with staff. Another could prepare a repair-ready plan for a person to approve. A third could make bounded commitments automatically, returning unfamiliar cases for investigation. These directions differ in the behavior they enable, the information they require, and the cost of an incorrect decision. Choosing among them is a substantive product decision; choosing whether the assistant appears in a sidebar or a full-screen chat comes later.

Here, AI can help make the alternatives concrete before a roadmap fixes one of them. A product manager could build a thin version of each interaction, connect it to representative data in a safe environment, and watch how staff and customers interpret the result. One version might expose missing inventory information. Another might show that customers will accept uncertainty if the next step is clear. Those findings can change both the scope and the interface.

The scarce resource may now be the opportunity to learn from real users rather than the ability to produce another mockup. Several beautifully rendered versions can consume that attention while testing the same premise. A useful comparison should make a different explanation or mechanism confront experience. For example, does confirming the part before scheduling create more value than making the initial booking faster? Under what conditions would the additional wait cause customers to leave?

This is also where prioritization should become more ambitious. Lower construction cost can make it sensible to examine an option previously rejected as too expensive to explore. It does not make operating cost, integration difficulty, or the consequences of a broken promise disappear. A small first release should resolve a consequential uncertainty in the larger vision while still delivering a coherent experience to the people using it.

The future product becomes credible through that sequence. Reliable intake could make better planning possible; reliable planning could support limited delegation; evidence from completed repairs could reveal where further autonomy helps. Each expansion should depend on a capability established by the previous work. A list of increasingly powerful features would not explain why the next one should succeed.

3. The specification now reaches into behavior

Traditional product work often separated an account of what should happen from the engineering work required to make it happen. The separation could be productive when each side supplied expertise the other lacked. It could also allow an important ambiguity to survive until implementation made a decision that the requirements had never confronted.

AI-assisted construction brings that confrontation earlier. A product manager can inspect a working flow and discover that “confirm the appointment” has several meanings. The interface might display a confirmation, the calendar might reserve a slot, and the inventory system might still have made no commitment about the part. Those are different states. The customer’s understanding depends on which one has actually occurred.

Suppose two customers request the last available visit window at nearly the same time. A prototype can appear successful when tested one conversation at a time and fail as soon as the requests overlap. The product needs a way to reserve the slot, identify when that reservation expires, and release it if the customer does not continue. If a network interruption causes an agent to retry an action, the system must distinguish a repeated request from a second booking.

These details belong in the product conversation because they determine the promise. A product manager does not have to implement every concurrency mechanism, but should be able to explain what must remain true and what the user experiences when the system cannot complete an action. The specialist engineering work becomes more focused when it begins with those distinctions already visible.

There is a useful boundary between exploratory building and a service people depend on. A disposable prototype can help test language, navigation, or whether a person understands a proposed interaction. Once the work changes persistent records, spends money, or makes commitments to someone else, the team needs a more exact specification of behavior, failures, and recovery. A sequence of prompts in one person’s chat history is a poor basis for another person to maintain the service.

AI can help turn that specification into tests. Given an available slot and a valid request, the service should reserve one appointment. Given an expired reservation, it should obtain a fresh confirmation before promising the visit. Given a missing part, it should follow the agreed path for uncertainty. These examples are useful because they expose decisions the system must make; their value does not depend on the format in which they are written.

The same scrutiny applies to requirements that are less visible in a demo. How long can a customer wait for a response? Which parts of the service must remain available when a model provider fails? How much information should the agent retrieve, and who may see it? What cost per completed repair can the business support? The answers shape architecture, user experience, and commercial viability together.

A more capable product manager should be able to use AI to investigate those trade-offs, construct a bounded implementation, and identify where specialist review is still required. This raises the standard beyond “I can generate an app.” It asks whether the person understands the app well enough to expose its consequential assumptions and change it deliberately.

Research has long cautioned that AI performance varies across tasks that look similar to a user. The Jagged Technological Frontier study tested 758 consultants on knowledge-work tasks and found substantial gains on work within the model’s capabilities, while emphasizing that the boundary was uneven.2 A product manager working with changing models needs task-specific evidence. Fluency at one part of the workflow should prompt investigation of the next part, rather than stand in for it.

4. The unit of success moves beyond the answer

The repair product could produce an excellent conversation and leave the customer with an unusable appointment. It could also give a brief, unremarkable response and enable a successful repair. A measurement system centered on how well the assistant speaks would miss the distinction that determines the product’s value.

I would start with the completed customer job, then work backward through the conditions that make it possible. For this product, successful repairs matter, alongside unnecessary visits, repeat failures, cost, and the customer’s ability to understand and control the commitment. Some outcomes will be difficult to observe directly, so the team may initially rely on operational signals. It should know which part of the outcome each signal represents and where the inference can break.

The design of τ-bench provides a useful research example. It evaluates an agent’s interaction with a simulated user, tools, and domain rules by comparing the resulting database state with the intended goal, and examines consistency over repeated trials.3 For our hypothetical product, an appointment written to the wrong customer record cannot become successful because the conversation sounds complete. Nor does a valid database entry establish that the technician ultimately repaired the appliance. Different stages require different evidence.

The analytical work becomes more demanding as the agent gains discretion. Suppose the service begins declining uncertain requests and its repair-success rate rises. It may genuinely be preventing wasted visits, or it may be excluding customers the business could have served with a different process. The team needs to examine the denominator, the absolute number of successful repairs, and the people who no longer receive a service. A higher rate can be a good result, but it must mean what the product goal says it means.

I would separate three questions when a promising system fails to improve the live result. Did users receive the intended behavior? Was that behavior directed at the right outcome? And did the measurement give us a valid comparison? These questions lead to different investigations.

First, examine what was actually delivered. If an inventory connection regularly fails and the service falls back to ordinary booking, the experiment may mostly measure the old workflow. Logs need to show which version acted, what information was available, where the fallback occurred, and what the customer received. The claim “the new model is live” is too coarse to diagnose this.

Second, examine the objective. A model rewarded for completing bookings may push ahead with insufficient information. A product goal centered on avoiding failed visits may over-encourage refusal. The task is to construct an objective and constraints that express the intended trade-off, then inspect where the actual behavior departs from it. Optimizing the wrong measure more effectively can make the product worse.

Third, examine the comparison. Customers reaching an automated channel may differ from those calling an employee. Their requests may be easier, or the initial rollout may occur at businesses with unusually clean data. Before attributing a result to the system, we need to account for those differences and establish the incremental effect of the intervention. Randomized or otherwise credible comparisons can help where practical, with the limits of the estimate remaining visible.

AI can contribute materially to each investigation. It can help formulate competing explanations, write and inspect queries, trace cases through a workflow, and suggest the segmentation most likely to distinguish a failure of delivery from a failure of the product idea. But a generated query can reproduce the same mistaken definition as the dashboard it is meant to check. The analyst needs to verify the event semantics and source data, not merely the syntax.

There is a further effect over time. If the service favors easily diagnosed repairs, businesses may specialize in those jobs, and the data available to improve the system may become less representative of other needs. A policy changes behavior, which changes the information used to justify the next policy. Long-term observation and deliberate exploration matter because early performance cannot describe every consequence of scaling the product.

This is where analytical thinking should grow with AI. The expanded capability is the ability to examine more of the causal chain, including the parts that are inconvenient for the preferred story. It should make a product manager harder to mislead with an attractive aggregate, whether that aggregate was produced by a colleague or a model.

5. More alternatives require better taste

A lower cost of generating solutions does not automatically create a broader search for the right one. Many variations can share the same assumption about the user, and a persuasive explanation can make that assumption harder to notice.

In a short-story experiment, Anil Doshi and Oliver Hauser found that AI-provided ideas improved individual creative output while making the resulting stories more similar to one another. Subsequent research by Yun Wan and Yoram Kalman found that more diverse AI-generated inputs could preserve diversity in a related writing task.45 The design of the collaboration mattered. For product work, I take that as a reason to examine how alternatives are generated, rather than count them and assume the search has become more ambitious.

The repair example could yield fifty versions of an assistant that books appointments. A genuinely different alternative might help technicians resolve simple faults remotely, help businesses predict parts needs, or help customers decide whether a repair is worth attempting. Each expresses a different theory of where value is being lost. The product manager needs to compare those theories, including the reason the company could deliver one better than its alternatives in the market.

This is a useful meaning of product taste. It includes recognizing which problem deserves a coherent experience, what information a person needs at the moment of choice, and which complexity the product should absorb rather than transfer to them. A chat interface might help someone explain an unfamiliar fault. A clear calendar with a few meaningful choices might be better for selecting a time. The product should not require conversation merely because the underlying model is conversational.

The commercial model is part of the experience as well. If a platform earns more from completed bookings than from successful repairs, its incentives can conflict with both customers and service businesses. A cheap AI interaction does not repair that relationship. A product manager should examine who receives the benefit, who incurs the cost, and what each participant is encouraged to do as volume grows.

AI makes it feasible to test a broader set of economic and behavioral possibilities before committing. It can help compare policies, identify where a business rule produces a perverse incentive, and generate cases that expose the weakness. Those simulations are useful for discovering what to investigate. The next step must still connect them to how actual participants respond.

Internal challenge needs similar care. Asking a second model to criticize a proposal can help, but its value depends on whether it introduces a stronger explanation, a relevant source, or a test that could defeat the proposal. If the critic sees only the polished summary, it may never encounter the assumption most in need of challenge. If it is rewarded for reaching agreement, another round may produce a more elaborate consensus without improving the underlying account.

Li Zhang and Kevin Ashley’s reflective legal-argument work makes one practical distinction explicit by separating scrutiny of factual support from rhetorical refinement. The method improved grounding and appropriate abstention in its bounded task, which used predefined legal factors.6 In product work, a corresponding question is whether the evidence supports the proposal before we invest in making the proposal more persuasive. The system must also be allowed to discover that an important factor was never included.

The stopping decision belongs here too. More research, more agents, or a stronger model can become another way to avoid choosing. Once a further investigation is unlikely to change the decision enough to justify its cost and delay, the team should act within the uncertainty it has identified. Conversely, a costly commitment should not proceed merely because the first working prototype makes its future feel inevitable.

6. A different standard for product management

The traditional product manager was expected to connect customer needs, business priorities, and the capabilities of a team. That responsibility survives. What changes is the amount of ambiguity they should be able to remove before asking others to commit, the distance they can carry a useful idea themselves, and the depth at which they must understand a system acting on the customer’s behalf.

A useful comparison therefore concerns the work product, rather than a list of new tools. A research summary can become an evidence-backed model of the problem with competing explanations still visible. A requirements document can be accompanied by a working interaction and tests exposing the disputed behavior. A launch report can reconstruct which treatment users actually received and whether it changed a meaningful outcome. Each artifact should help the team make a decision it could not make as well before.

None of these forms is mandatory for every problem. An experienced product manager may establish that a simple policy correction is more valuable than another prototype. They may find that the next uncertainty can only be resolved by speaking with a customer or obtaining data the organization does not have. Technical fluency should improve that choice, not create pressure to demonstrate code in situations where code supplies little information.

The standard should nevertheless be concrete. Give a product manager an unfamiliar problem and ask them to explain whom it affects, what evidence supports their interpretation, and why a proposed intervention could change the outcome. Then ask them to make the experience inspectable, identify a consequential failure, and revise the design when a material assumption changes. AI can assist throughout. The person should be able to account for the choices, locate the supporting evidence, and recognize where the investigation remains incomplete.

For a senior product leader, the same standard extends to the allocation of effort. They should be able to distinguish a missing feature from a platform constraint, explain when a local improvement creates costs elsewhere, and judge whether the next investment should deepen an existing capability or pursue a different opportunity. A rewrite may be justified if it removes a structural obstacle to the vision; a refactor may preserve more value if migration would consume the team before the customer benefits. The answer should follow the economics and dependencies of the work.

This does not require a product manager to become the most knowledgeable person in every discipline. It requires them to work productively across the boundaries that still matter. Engineering expertise can reshape the vision before it becomes a promise. Design research can reveal that an apparently efficient workflow asks users to make decisions they cannot understand. An operations specialist can identify the exception that determines whether the service works outside a demonstration. AI should make those contributions easier to examine and incorporate.

There is a developmental consequence. As some tasks become easier to produce, the old signals of seniority become less informative. A polished strategy document may reveal little about the author’s ability to handle an unfamiliar trade-off. Leaders should examine how someone responds to a contradiction, what they choose to investigate, and whether they can turn a correction into a better product decision. The purpose is to develop a person whose useful scope grows with the tools.

7. When AI can do more of the thinking

The more difficult future begins when AI becomes good at work we currently regard as the product manager’s distinctive contribution. It may identify unmet needs, formulate better strategies, design experiments, and discover why a seemingly successful product is harming the broader business. We should be willing to consider that possibility rather than define the profession around a permanent exclusion from machine capability.

Imagine a system that detects a recurring failure across repair jobs, proposes a new diagnostic interaction, constructs a prototype, and tests its behavior against prior cases. It then identifies the uncertainty requiring contact with customers and prepares a bounded experiment. After the experiment, it evaluates whether the result warrants a larger commitment, taking account of cost, service quality, and the business’s capacity to fulfil it.

That would move the product manager’s work again. Some people might direct a much larger portfolio of investigations. Others might specialize in establishing the right objectives, developing evaluation methods, or working through cases in which the system’s representations fail. Some existing responsibilities could disappear. The allocation should depend on demonstrated capability and the purposes of the organization, not a promise that every current job description will remain necessary.

A more capable system would also challenge our understanding of the product itself. Instead of releasing a fixed interaction for every customer, a service could adapt how it obtains information and arranges work while preserving the commitments that define its value. The product manager would need to understand which variations are useful, which can be evaluated, and which would change the promise people agreed to rely on. AI could help answer those questions too.

The opportunity is much larger than making today’s feature pipeline cheaper. A small repair business might offer customers the coordination and attention once available only through a large operation. A product team might examine an underserved segment whose needs previously looked too specific to justify the research. More people could turn a well-understood problem into a service without first obtaining permission from a large organization to explore it.

There are still limits that capability alone does not remove. Physical work takes time, customers have preferences that cannot simply be optimized away, and resources allocated to one opportunity are unavailable for another. More intelligence can improve how we understand those trade-offs. It can also reveal opportunities hidden by the effort it once took to examine them.

Louise asks for access to the encounter because the recording cannot settle the question she has been asked. Product management should retain that instinct as its means of investigation become more powerful. Use the model to reach farther into the problem, make the proposed experience real enough to challenge, and stay with the work until the evidence shows what it has changed. The profession’s future will be shaped by the possibilities that discipline allows us to build.


Notes

Footnotes

  1. Eric Heisserer, Arrival, screenplay, printed pp. 11–13. The quoted words appear on printed p. 13 in the available revised screenplay. Louise’s interaction with Colonel Weber supplies the opening example. Screenplay. Back to reference 1

  2. Fabrizio Dell’Acqua and colleagues, Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality (2023). Harvard Business School research account and paper link. The study used 758 consultants and then-current models; its task-dependent findings do not locate the capability boundary of a current deployment. Back to reference 2

  3. Shunyu Yao, Noah Shinn, Pedram Razavi, and Karthik Narasimhan, τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains (2024). The benchmark uses simulated users, domain tools and policies, final database state, and repeated-trial consistency. The essay borrows its evaluation distinction without treating the benchmark as an estimate of present-day repair-service performance. Back to reference 3

  4. Anil R. Doshi and Oliver P. Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances 10 (2024). The experiment concerns short-story writing, not product-strategy outcomes. Back to reference 4

  5. Yun Wan and Yoram M. Kalman, Diverse AI Personas Can Mitigate the Homogenization Effect in Human-AI Collaborative Ideation, revised 16 March 2026; published in Computers in Human Behavior: Artificial Humans. The study varies AI plot inputs in a writing task; it does not establish independence among product-review agents. Back to reference 5

  6. Li Zhang and Kevin D. Ashley, Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation (2025). The study uses a predefined factor representation and bounded reflection. Applying the distinction between support checking and rhetorical refinement to product proposals is an extension developed here. Back to reference 6

Back to top

IN THIS SERIES

Organizations and products in the age of AI

  1. When capability outruns the organization
  2. Product management with the means to build — You are here

THE INDEX

Find a thread.

Loading the index…