When the framework is the problem

In this post

In Dune, Paul Atreides asks why the machines harvesting spice are not protected by shields. It is a reasonable question from someone accustomed to relying on that technology. The answer depends on something his experience has not prepared him to consider. “A shield’s a death sentence in the desert,” Kynes explains, because it attracts the worms.1 The equipment behaves as designed; the environment changes what using it will accomplish.

A difficult problem often arrives with a familiar name. It is an approval problem, a research problem, an adoption problem. The name helps us organize the work, but it also imports assumptions about what matters and which solutions deserve attention. We can spend considerable effort improving an answer before discovering that the question excluded the most important part of the problem.

This is where I think the habits of a legal mind and a product mind become particularly useful together. The discipline I would take from litigation examines what entitles us to reach a conclusion, which distinctions change its meaning, and whether it survives the strongest supported objection. Product thinking takes that understanding into the design of something people can use, asking what would change in their lives, how the system would behave, and whether the improvement would remain valuable at scale. Together, they offer a way to pursue unfamiliar possibilities without treating novelty as permission to abandon rigor.

We should use AI to expand our capacity for this kind of invention, including the ability to discover that an entirely different product is possible once the original framing is questioned. The ambition should extend beyond accelerating the work we already know how to assign.

First principles have to constrain the argument

The beginning of first-principles thinking is deciding which premises deserve to survive. A useful distinction is between what the evidence establishes, what we infer from it, and what we have assumed because it makes the problem easier to work with. Once these are collapsed into a single account, every subsequent step can look reasonable while depending on something no one has actually established.

The comparative habit of legal reasoning is useful here. Two situations can share a vocabulary and still operate differently because the relevant obligations, incentives, or conditions differ. Before transferring a conclusion from one to the other, we should identify which features made it valid in the first place. Similarity becomes a hypothesis to examine rather than sufficient reason to reuse the answer.

Consider a hypothetical team asked to automate an approval process. One approval establishes that money is available, another that a technical risk is acceptable, and a third that someone has authority to commit the organization. Combining these into one automated decision because they share a label could eliminate the distinctions that determine what the system may do. Equally, preserving every existing handoff could retain steps whose only purpose was to move information that the product can now obtain directly.

The useful inquiry works in both directions. We ask what each step protects, but also whether its current form is necessary to provide that protection. A requirement to stay within a budget does not, by itself, establish that every request needs a manager to recheck the balance. Conversely, a manager’s approval of one request does not establish permission for an agent to make every similar commitment thereafter.

This changes the design problem. Instead of reproducing a sequence of approvals more quickly, the team can investigate a service that knows which conditions a request must satisfy, obtains the necessary evidence, and allows it to proceed within explicitly delegated limits. Some checks might run together; some handoffs might disappear; genuinely novel risks might still require specialist attention. The route to greater autonomy comes from understanding the decisions well enough to separate them.

That is why rigorous first-principles thinking can support a larger ambition. It identifies constraints we must respect while exposing arrangements that can be redesigned. The proposed model has to explain the difficult examples as well as the convenient ones, and it must be specific enough to guide an implementation. Otherwise, we have replaced an inherited assumption with a more attractive assertion.

The challenger inside the argument

An unfamiliar idea needs a serious attempt to establish where it would cease to be defensible. This is particularly important when the idea is promising enough that everyone involved would like it to work.

I think of an internal adversarial challenger as a discipline of reasoning, whether performed by another person or deliberately undertaken ourselves. The challenger should identify the premise on which a proposal depends, construct the strongest supported alternative, and specify what evidence would distinguish them. It should be possible for the challenge to defeat the proposal. Producing objections that its author can comfortably answer adds very little.

Return to the approval example. Suppose two requests each fit within the available budget when assessed separately, but together exceed it. This objection exposes an assumption about independence. Answering it could require a mechanism that reserves capacity when a commitment is made, rather than a more accurate recommendation about whether each request looks acceptable. The challenge has identified something the product needs in order to exercise more autonomy reliably.

Limiting an idea is therefore part of developing it. A boundary condition can reveal a narrower launch population, a missing source of evidence, or a more ambitious architecture than the initial feature proposal allowed for. A blanket rejection would miss those possibilities, while a determination to defend the original design would leave its weakness intact.

The inquiry must also escape circular reasoning. A forecast may assume adoption because the proposed roadmap is valuable, while the roadmap is justified by that same forecast. Each document appears to support the other without establishing why a customer would change their behavior. A useful challenge sends the team outside that circle, perhaps to observe the existing workaround, test whether the proposed capability changes a decision, or establish what the customer would give up to use it.

Internal debate cannot settle those questions on its own. We should borrow the rigor of adversarial examination without turning colleagues into opponents or mistaking a successful presentation for validation in the market. The measure of a challenge is what it helps us discover and change, not how difficult it makes the meeting.

What AI should help us examine

These distinctions suggest a more demanding role for AI than generating a proposal and asking another agent whether it is convincing.

If several agents are asked to optimize an approval queue, they may offer competing answers while all accepting that the queue is the correct unit of work. Assigning different personas does not ensure that any of them will ask which decisions the queue contains or whether those decisions need to remain sequential. A system can explore many alternatives inside a frame without testing the frame itself.

I would give that earlier inquiry an explicit place in the work. One line of investigation might reconstruct the requirements from source material; another might look for examples that contradict the proposed classification; a third might develop an alternative design that satisfies the same underlying obligations. Their value would depend on the evidence and reasoning they add, not the number of agents involved. They should be able to obtain missing information and leave a consequential disagreement unresolved when the available material cannot settle it.

There is a useful, bounded research example in Li Zhang and Kevin Ashley’s work on reflective legal argument generation. Their system separates analysis of factual support from rhetorical polishing and tests whether an argument should be abandoned when the supplied factors cannot support it. The study reports improvements in grounding and appropriate abstention, but uses predefined legal factors. Identifying the right factors in unfamiliar material remains a further problem.2

That is where I would place a more ambitious product test. Give the system a problem without naming the decisive distinction, then assess whether it discovers the distinction, supports it, and applies it within defensible limits. Change a material fact and examine whether the analysis changes appropriately; change only the wording and examine whether the reasoning holds. A benchmark that supplies the disputed classification has already performed part of the work we wanted to evaluate.

More capable systems could eventually help formulate these tests, discover alternatives we had not considered, and propose changes to the evaluation itself. I would welcome that progression. We should still evaluate a changed method against criteria it has not unilaterally rewritten, because improvement becomes difficult to establish when the system can change both its answer and what counts as success.

The creative opportunity is substantial. An assistant that helps examine only the first proposed solution can make that solution better. A system that helps investigate several defensible formulations of the problem could reveal opportunities we would never have reached by improving the first one. Whether it does so is something to test, rather than an automatic consequence of adding a reflection loop.

Product management after the handoffs

The same scrutiny should reach the organization building the product. If we are willing to question whether a familiar workflow still fits its purpose, we should also ask whether familiar divisions between product management, design, engineering, and domain expertise still fit the capabilities available to us.

One plausible response is to let individuals carry more of a product from idea through delivery; another is to retain specialist roles while changing how they collaborate. The distinction turns on which expertise an AI-assisted team can actually exercise and which it still needs to obtain elsewhere. A product manager completing a prototype does not establish that specialist engineering is unnecessary; equally, a specialist’s importance in one difficult context does not establish that every existing handoff is essential.

There is empirical support for investigating those boundaries. In the published Cybernetic Teammate field experiment, involving 791 professionals working on product innovation challenges, individuals using AI matched the performance of teams without it, and AI helped participants produce proposals combining technical and commercial perspectives. These were results about innovation tasks, not proof that one person could replace an entire product team through implementation and operation.3

For the present transition, I would organize more work around small groups capable of taking a consequential question through investigation, a working intervention, and evaluation. Product leaders should examine implementation assumptions; engineers should participate in deciding whether an intervention addresses the user’s problem; domain specialists should help define the tests before the product commits to an inadequate representation. Working together earlier gives a legal objection the chance to change the design, and a technical possibility the chance to change what the domain expert thought was necessary.

This is the convergence worth pursuing. Expertise remains deep enough to recognize a consequential distinction, but its influence extends into what gets built. A product manager should be able to carry a disputed assumption into a prototype and an instrumented test, with specialist help where the problem requires it. A legal specialist should be able to explain what would need to change for a proposed service to work, rather than be consulted only after its important choices have been made.

For a more mature AI environment, we need a less comfortable assumption. Models may become good at identifying bad hypotheses, proposing better architectures, and challenging expert interpretations. Reserving strategic reasoning permanently for people while delegating only production to machines would build an untested limit on AI capability into the organization itself.

Under that scenario, product leadership would increasingly concern the design of a decision-making system. Which problems deserve resources? How should competing explanations be examined, and what evidence permits a consequential action? AI could contribute substantially to those analyses. The organization would still need to assign responsibility for its commitments and provide a way for affected people to challenge them. The question becomes how to combine capabilities and accountability as both the work and the technology change.

What we should change now

We can begin without waiting for that future. On an unfamiliar initiative, I would ask the team to make the assumption behind its preferred approach explicit, then give someone outside the proposal’s authorship access to the underlying evidence and responsibility for testing the strongest alternative. Their output should identify a consequential disagreement and a way to resolve it. The person accountable for the decision must then explain whether it changes the plan, calls for an experiment, or remains a risk worth accepting.

We should fund that investigation as part of delivery. A challenge discovered after every resource and deadline has become immovable is less likely to influence the work. When investigation reveals a better opportunity, leaders need to be willing to revise the objective and move the resources, including when the original plan was their own. Conversely, independent scrutiny should not become an unlimited veto. Its depth should reflect the consequences of error, the reversibility of the commitment, and whether another round of inquiry is likely to change the choice.

We should also develop people against the reasoning we expect them to exercise. Have them form an initial view, reconstruct the strongest opposing explanation, and show how a changed fact would affect their conclusion, using AI to extend the investigation rather than conceal where their understanding ends. Then ask them to turn the resulting insight into something inspectable. The discipline is incomplete until the revised understanding changes a prototype, an experiment, an operating mechanism, or a decision about what should be built.

There is a different standard for innovation in this approach. Originality has to survive scrutiny, while scrutiny has to remain capable of revealing something more ambitious than the existing plan. The legal mind helps expose what a conclusion depends on; the product mind uses that understanding to construct and test a better possibility. Neither contribution is complete in isolation.

We should use AI to make this kind of investigation accessible to more people and more teams. The most consequential result may be a product no one had put on the roadmap, because examining the assumptions revealed an opportunity the roadmap could not express. Leaders should make room for that discovery, give it a credible test, and be prepared to change the organization when the evidence warrants it.


Notes and sources

The approval workflows are hypothetical examples. The organizational recommendations and mature-AI discussion are proposals and conditional arguments, rather than reports of deployed capabilities.

Footnotes

  1. Jon Spaihts, Denis Villeneuve, and Eric Roth, Dune, final shooting screenplay, 19 June 2020, printed p. 56 (PDF p. 57), adapted from Frank Herbert’s novel. The quoted sentence is Kynes’s reply to Paul’s question about shielding the spice crawlers. Screenplay. Back to reference 1

  2. Li Zhang and Kevin D. Ashley, “Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation” (2025), arXiv:2506.02992v2. Paper · Full text. The experiment uses predefined factors for a three-part argument task. Section 7.1 identifies the limits of that representation, bounded reflection, and automated evaluation. The proposed product tests in this essay extend beyond what the study demonstrates. Back to reference 2

  3. Fabrizio Dell’Acqua and colleagues, “The Cybernetic Teammate: A Field Experiment on Generative AI and Teamwork,” Organization Science, published online 12 June 2026. Published paper. The preregistered experiment studies product innovation tasks with 791 professionals. Its findings support investigation of cross-functional collaboration; they do not establish that the proposed organizational model works across entire product lifecycles. Back to reference 3

Back to top

THE INDEX

Find a thread.

Loading the index…