What would it take to build an autonomous advocate?
Cognition describes Devin as a software engineer that can plan, write, test, and ship code. Meta’s Muse can negotiate on a person’s behalf and continue working after they close the app. OpenAI’s dots are designed to remember context, work across connected applications, and decide what to do next between conversations. These products invite us to delegate work that previously required our continuing direction, while retaining control over the decisions we want to make ourselves.1
Are we on the cusp of extending the same to litigation work? How far could that arrangement extend into litigation? An assistant that investigates a dispute, develops the arguments, and handles the proceedings would, arguably, be enormously valuable. For an experienced lawyer, it could expand the work a small team can undertake. For someone who cannot afford representation, it could change whether they are able to pursue a right at all.
I expect that at one point in time in the future, autonomous advocacy will eventually become possible; but not yet. With today’s systems, including the new personal agents, such cannot yet reliably sustain the combination of judgment it requires which a human can. A case develops over time, often going through the motions as an opponent’s attempts to defeat it, and the information that changes its direction may initially look less important. The system must recognize that change, decide which skills to apply, and preserve the consequences of earlier decisions while responding to new ones.
The gap is between performing legal tasks and directing them through a changing case. Understanding that gap should help us decide what to build, what to test, and how much useful assistance can reach people before the entire job becomes safely delegable.
A different kind of correctness
Software’s mathematical foundations give an engineering agent a valuable source of feedback. It can execute a change, observe its behaviour, and repeat tests in an isolated environment. Requirements and design still require judgment; Cognition’s FrontierCode evaluation considers whether maintainers would accept the work, beyond whether it passes tests.2
In advocacy, the parties may dispute the principle by which the result should be judged. They may agree on what happened while disagreeing about its legal significance, or agree on the law while contesting whether the evidence establishes its requirements, or disagree on what happened altogether. A principle that appears to support the client must withstand a judge asking why it should apply here, where its limits lie, and whether those limits are consistent with the result sought.
An advocate may respond by narrowing the principle, but discover that the narrower version no longer supports the claim. Alternatively, the law may favour the opponent while their evidence fails to establish the necessary facts. Recognizing which problem is open to challenge determines the work that follows. More research into a legal proposition already conceded would consume time that should have gone into investigating the evidence.
Research illustrates how a system can possess information without making that connection. In DLawBench’s simulated consultations, one model gathered most of the relevant facts about a disputed bonus but developed a wage-underpayment strategy before adequately establishing whether the bonus was owed and which employer was responsible. The account was largely faithful to the conversation; its organization around the wrong legal question made the advice unreliable.3
The example matters because a better memory of the conversation would not, by itself, have repaired the analysis. The system needed to reconsider what had to be established before the proposed claim could succeed. An advocate directing the AI can supply that correction. An autonomous advocate has to recognize the need for it.
The case changes what matters
That recognition must continue long after the first consultation. A document received months earlier may become decisive when a witness explains who made a decision. An alternative advanced on an assumed fact may become the primary case after new evidence arrives. The system needs to preserve the difference between an allegation, an admission, and a temporary assumption while continually reassessing their significance.
The new personal agents already address persistence. OpenAI describes dots as remembering context across ongoing work, and Meta explicitly identifies long context, extended instruction-following, and multi-agent coordination as areas of Muse’s training.4 The difficulty therefore cannot be dismissed as a chat window forgetting yesterday’s conversation.
The harder problem is knowing what to recover and what to reconsider. If an agent searches the file using an outdated understanding of the dispute, it may find excellent support for a theory it should now question. A summary can accurately preserve what seemed important when it was written while leaving out the detail that later becomes consequential. Returning to the original material helps only if the system recognizes a reason to return.
Earlier long-context research found that models struggled increasingly as inputs grew when finding the relevant passage required connections beyond matching words.5 The study did not evaluate today’s personal agents. It identifies a difficulty that becomes particularly consequential when the relationship making a fact relevant changes as the legal argument develops.
This becomes especially demanding during a hearing. A question from the bench may require counsel to interpret the concern, retrieve the relevant material, reconsider a legal distinction, and assess what the answer concedes elsewhere. Preparation permits those inquiries to proceed in stages. Live advocacy often requires them together, under the pressure of an exchange whose direction counsel does not control. Sometimes the right response is to seek time or instructions; recognizing that need is part of the judgment too.
Who directs the work?
A natural response is to divide the work among specialist agents. One investigates the facts, another researches the law, and others challenge the argument or prepare the draft. That could substantially expand the investigation, but it also creates a question about who decides when the assignments themselves are wrong.
Suppose the coordinator treats a dispute as turning on the interpretation of a contract. Its researchers might find accurate authorities and its drafters produce persuasive submissions, while the decisive obstacle concerns whether the client can establish the facts needed to invoke that interpretation. Each specialist can perform its assigned task competently while the system as a whole pursues the wrong route.
The coordinator must allow a specialist’s finding to reopen the plan and determine how that changes work already completed. Studies of multi-agent systems have identified failures in shared understanding, coordination, and verification.6 An autonomous advocate needs those relationships to remain reliable even when the next necessary action was never part of the plan.
LegalWorld investigates these connections through simulated consultation, drafting, and courtroom work, with earlier decisions affecting later stages. The evaluated models performed more strongly in formal drafting than in the courtroom capabilities measured. The simulation still simplifies significant procedural branches and settlement, leaving important parts of autonomous representation to be tested.7
Permission controls address another part of the problem. Muse’s separate Sentinel governs external actions and seeks approval where required.4 Such protections can prevent an impermissible operation without establishing whether a permitted filing advances a sound case. A client may approve an argument precisely because they lack the knowledge to recognize its weakness. A safe service needs to distinguish permission to act from the competence needed to advise the person granting it.
This is why the lawyer’s continuing contribution needs to remain visible. If counsel identifies each decisive issue, chooses the next investigation, and repairs the connections between outputs, that directing judgment is part of what makes the system useful. Removing counsel would require replacing that contribution, not simply allowing the same tools to run longer.
What would change the assessment?
A serious test should let the case develop without continually telling the agent what has become important. Information should arrive in stages, an opponent should be able to adapt, and new evidence should sometimes require abandoning a previously plausible route. The evaluation should examine whether the agent discovers the relevant change, revises the affected arguments, and chooses a defensible next step within its mandate.
Record the human interventions required to keep the case on track and the time spent checking or repairing the work. Handling more of the matter while requiring fewer consequential corrections would support greater delegation. Where several defensible approaches exist, independent practitioners should assess the reasoning and consequences rather than require agreement with one model answer.
Learning from the result is equally important and equally difficult. A victory may rest on a consideration the agent scarcely investigated. A loss may reflect an error by the court, an unavoidable evidential weakness, or a poor decision during preparation. Teaching the system to repeat everything associated with winning would confuse those possibilities. We need to compare the reasons given with what was anticipated, what was reasonably knowable, and what the advocate could have done differently.
LegalWorld’s exploratory work found that reflection on completed simulations improved later performance on related benchmark cases.7 For real advocacy, a lesson must retain the circumstances that made it useful and remain open to revision. A system that turns a case-specific result into a rule for the next client could become more confident by learning the wrong thing.
Directing the inquiry
Treat AI as additional capacity for an investigation that the advocate continues to direct. Let it develop competing routes far enough to expose what each requires, then ask it to reconstruct the strongest answer the opponent could give. Counsel decides which uncertainty deserves the next hour and what its resolution would change.
That judgment needs to move with the case. New evidence may require reopening an earlier interpretation; an adverse legal conclusion may redirect the inquiry towards the facts needed to apply it. Keeping those connections intact is part of the work we retain when delegating individual tasks. The question to carry into practice is whether AI has made us better able to discover where our case needs to change, as well as better equipped to advance it.
Notes
Footnotes
-
Provider descriptions checked 2 October 2026. Cognition on Devin, OpenAI, Getting started with your dot, and Meta, Introducing Muse. These describe capabilities and design aims, not measured performance in autonomous litigation. The assessment of readiness in this post is the author’s position, not a report of a direct comparative trial of these products. Back to reference 1
-
Cognition, Introducing FrontierCode, particularly “Beyond Unit Tests.” Software evaluation also involves judgment about design, requirements, scope, and risk; the comparison concerns the forms of feedback available, not a claim that every software problem is formally decidable or easily reversible. Back to reference 2
-
DLawBench: Evaluating LLMs Through Multi-Turn Legal Consultation, §4 and Appendix B.1, “High coverage, wrong legal route.” The example is an evaluated simulated consultation. The researchers’ reference analysis concerns bonus entitlement, the meaning of payroll information, and employing-entity identity. The benchmark is not a study of real client outcomes or every current agent. Back to reference 3
-
OpenAI’s dot documentation; Meta, How We Built Safety Into Muse, opening discussion and “A Built-In Sentinel.” Published protections should not be mistaken for an absence of safeguards, nor do they independently establish litigation competence. Back to reference 4 Back to reference 4-2
-
Modarressi et al., NoLiMa: Long-Context Evaluation Beyond Literal Matching, ICML 2025. The study investigates retrieval requiring associations without literal wording overlap in the models tested. The application to changing legal significance is an inference; the study did not evaluate dot or Muse. Back to reference 5
-
Cemri et al., Why Do Multi-Agent LLM Systems Fail?, 2025. The taxonomy addresses system-design, inter-agent, and verification failures across the examined frameworks and tasks. It does not establish a fixed limit on all multi-agent systems or directly test autonomous advocacy. Back to reference 6
-
LegalWorld: A Life-Cycle Interactive Environment for Legal Agents, §§4.5–4.6 and “Limitations,” June 2026. The setting is simulated Chinese civil litigation with model-based capability scoring and human validation of aspects of the environment. The reflection experiment concerns later same-cause benchmark cases. Results do not establish actual legal-service outcomes, performance across all jurisdictions, or live-hearing reliability. Back to reference 7 Back to reference 7-2