The apprenticeship of legal AI

In this essay

The opening epigraph of Frank Herbert’s Dune offers a warning about beginnings.

“A beginning is the time for taking the most delicate care that the balances are correct.”1

A learning system creates a new beginning whenever it turns an experience into a method. A correction that once affected a paragraph can begin to influence how future questions are framed, which sources are consulted, and what the system accepts as adequate support. An error can acquire a much longer life in this form. So can a useful insight.

This is what interests me about recursive agentic learning in law. Imagine a lawyer correcting an argument because the evidence supports a narrower proposition than the draft asserts. The immediate benefit is a better argument. A more ambitious system would also investigate the mistake, formulate a method for avoiding it, and test whether that method helps on work it has never seen. Later experience could refine the method further. A successful correction could improve work beyond the matter in which it was made.

Research on agents that extract and revise reusable guidance gives us reason to take this possibility seriously. It also exposes a difficult question. When a system learns from legal work, what should count as evidence that it has become better at it?23

Acceptance of a draft, success in persuading another agent, and a favorable court result answer different questions. None, by itself, establishes what the system should do differently next time. I think the prospects for learning legal systems depend heavily on whether we preserve those distinctions as the technology becomes more capable.

1. What an experience should teach

Consider a hypothetical dispute over a notice. A draft says that the recipient knew about a defect because an email was sent. Counsel corrects it. The available record establishes dispatch, but does not establish that the recipient read the message.

Several lessons could be extracted from this correction, with very different consequences. The system could remember that this recipient lacked knowledge, although that would overstate what counsel established. It could learn never to infer awareness from an email, which would be wrong in a later case containing an acknowledgment. Or it could require proof of personal awareness in every notice dispute, even where the governing provision makes delivery sufficient.

The useful lesson begins one step earlier. Identify the proposition that matters under the applicable rule, then examine what the evidence establishes about it. Dispatch, delivery, and awareness may need separate treatment. Sometimes only one of them matters. Sometimes the connection between them is a reasonable inference that should be defended rather than presented as an observed fact.

Notice how much of the legal work lies in deciding which question to ask. A general instruction to demand more evidence could make an argument less accurate by introducing a requirement the governing rule does not impose. A general instruction to be more cautious could suppress a sound inference. The conditions surrounding a correction are part of its meaning.

Suppose the applicable procedure also requires the sender to establish delivery. A gap in the evidence could then determine the issue without establishing that delivery never occurred. A system needs to preserve the difference between uncertainty about an event and a decision about what follows from that uncertainty.

The same difficulty appears when learning from judgments. Suppose a court expressly assumes a factual allegation for the limited purpose of deciding an application. A system that stores the allegation as a judicial finding has changed its status. Reusing the summary may then spread the error into research that was never concerned with the original application. The words can be accurately quoted while the proposition attributed to them is wrong.

This is why I would distinguish a reusable method of reading an authority from a reusable statement of law. The method might ask what the court actually decided, which facts it accepted for that purpose, and which propositions the decision leaves unresolved. A statement of law needs its own support and limits. The fact that a method was useful in one jurisdiction does not make the rule applied there available everywhere else.

For learning systems, the unit of improvement should therefore be a change in behavior with identifiable conditions of use. “Be careful with notices” is too vague to test. “Always require proof that a notice was read” is specific but often the wrong instruction. The harder and more valuable achievement is learning when a distinction changes the analysis.

2. What is beginning to work

It helps to separate two loops that are often grouped together. In one, an agent revisits the current problem, perhaps finding another source or responding to a critic. That extra work is useful when it can resolve a live objection or change the evidential basis of the answer. In the other loop, evidence from completed work changes the methods available for later problems. Repeated deliberation alone does not create that durable change.

There are also different places where learning can occur. A system may revise written procedures supplied to an otherwise unchanged model. It may train a model that selects tools and methods. Or it may change components involved in the models’ internal computation. Research called Recursive Multi-Agent Systems, for example, trains connections that transmit and refine internal representations between agents. Its mechanism is different from having several assistants exchange another round of text.4

The more immediately legible form of learning is procedural. In Experiential Reflective Learning, agents derive conditional guidance from previous task attempts and retrieve relevant guidance for new tasks. On the study’s Gaia2 search and execution tasks, in a simulated mobile environment, success increased from 48.3% to 56.1%. The evaluation used separate data environments from those used to accumulate experience. Simply supplying earlier task trajectories as examples did not produce the same benefit.2

The result does not establish improved litigation performance. It does demonstrate a distinction worth investigating in law. What transfers may be an abstraction about how to approach the work, rather than a more detailed memory of having done similar work before.

WikiSkill develops this idea by separating execution records, accumulated knowledge, and the procedures used by an agent. A proposed procedure can be rejected after testing while knowledge of the attempt remains available for subsequent proposals. Its experiments cover mathematical reasoning, search, spreadsheets, document questions, and interactive tasks.3

That separation has a useful implication for professional learning. Discovering why an attempted improvement failed can be valuable even when the attempted improvement should never be used. A record of that failure can help a later investigation without acquiring the authority of an instruction.

A stronger form of recursion would allow improved methods to help investigate subsequent failures, devise more discriminating tests, and propose further improvements. The object of learning could eventually include parts of the learning process itself. That is a possibility to investigate, not a capability established for litigation by these studies. Their environments and feedback do important work. We still have to specify what the equivalent of successful task completion would be in legal practice.

3. The difficulty of rewarding judgment

A redline contains more information than a preference between two drafts. Moving a paragraph may improve the order of persuasion. Replacing a citation may correct the source of a legal proposition. Removing an assertion may reflect missing evidence, a strategic concession, or a change in the client’s instructions. Learning the right lesson requires distinguishing those reasons.

Even explicit approval is incomplete feedback. A lawyer might accept a draft after checking its authorities but before checking its account of the evidence. Another might approve an argument for a particular application while rejecting it as advice on the merits. Treating both approvals as endorsements of every reasoning step would give the feedback a scope it never had.

The correction may also introduce information the system never had. Suppose counsel fixes an inference by supplying an omitted document. Learning only from the difference between the drafts could misdiagnose a failure to obtain evidence as a failure to reason about it. The more useful change might concern when to retrieve a source or ask the lawyer for it.

A favorable disposition adds another problem. Suppose counsel advances three alternative grounds and succeeds on one. The result cannot validate all three. It may provide useful evidence about the ground addressed, but it still does not isolate the contribution of a particular passage, research sequence, or drafting technique. A learning algorithm needs a more precise target than “repeat what appeared in the winning document.”

SkillFlow addresses a related technical problem. It trains a supervisor over an executor and a changing skill library, using a flow-based objective to retain multiple high-reward strategies and derive signals about the importance of individual steps. Those signals help guide which skills to develop or remove.5 This is a meaningful advance in the machinery of credit assignment. It does not determine whether the reward represents sound legal analysis, or establish that a particular drafting move caused a court’s decision.

Persuasion makes the distinction especially important. In one study of trait-conditioned legal argumentation, advocates competed before a model acting as judge, using ten synthetic cases. The judge and advocates shared the same model backend. The authors explicitly acknowledge the possibility that the judge favors linguistic patterns like those it generates.6 An improvement in that environment may reflect a more effective argument, a better fit to the evaluator’s preferences, or some combination. Its meaning depends on further testing.

There is a more concrete starting point in research by Li Zhang and Kevin Ashley. Their reflective legal-argument system separates scrutiny of factual support from rhetorical polishing and tests whether agents abstain when the provided factors cannot support the requested argument. The study reports improvements in grounding and abstention, while acknowledging that its predefined factors leave much of the interpretive work outside the experiment.7

That limitation identifies an important next question. Can a system improve at finding the legally relevant distinction in a full record, rather than only using distinctions already supplied to it? Better legal reasoning would include identifying a stronger opposing argument, recognizing an exception, or discovering that the apparently decisive inconsistency disappears when the witnesses’ questions are read in context. Some of these improvements make the client’s preferred conclusion harder to maintain. A reward that penalizes that result would train the system away from the judgment we wanted.

4. Testing a change in reasoning

The practical response is to make the proposed lesson confront cases that could defeat it.

For the notice example, I would test more than whether the system stops equating dispatch with awareness. I would also ask whether it recognizes an acknowledgment, whether it applies a hypothetical provision requiring delivery rather than personal awareness, and whether it explains a defensible inference without turning it into certainty. A method that responds to every case by withholding a conclusion has learned avoidance.

The test should also distinguish sensitivity from instability. Changing the fact that matters should change the analysis. Changing names, formatting, or irrelevant background should not. Where reasonable interpretations remain open, the evaluation should examine whether the system identifies and supports them. Forcing a single preferred answer can reward conformity precisely where the task requires judgment.

Formal reasoning can help with part of this work. In a study of insurance coverage analysis, unguided conversion of contracts into logic programs produced worse results than answering directly with an LLM. Providing a structured framework improved the encodings.8 A successful logical derivation is only as useful as the representation from which it starts. An exception attached to the wrong clause can yield a perfectly consistent answer to the wrong version of the contract.

I would therefore test the interpretation as well as the subsequent reasoning. Does the procedure preserve the scope of a qualification? Does it distinguish an asserted fact from an accepted one? Does it recognize that a missing record leaves a proposition unresolved? These are narrower questions than whether the draft “looks more expert,” and they are more useful for diagnosing an improvement.

Independence matters at several levels. Evaluation examples should come from work separate from that used to create the lesson. A paraphrase of the originating matter is not an independent example merely because it has a different filename. The comparison should preserve the same available evidence and a specified resource budget, so that an alleged gain from learning is not simply a gain from receiving another document or more computation.

The evaluator also needs examination. Research on LLM judges for legal document recommendation measures agreement with human experts across different dimensions, rather than treating the evaluator’s verdict as self-validating.9 For verification, the second agent’s value lies in checks that expose errors the first could miss. Giving two agents different job titles does not establish that their judgments are independent.

A learning system could help create new tests, particularly by finding variations under which a proposed method breaks. It could also suggest changes to the evaluation criteria. But allowing the same update to rewrite both the method and the criteria for accepting it would make improvement difficult to establish. Proposed changes to the test need evidence and review of their own.

This need not mean asking a partner to review every experimental variation. Mechanical checks, calibrated evaluators, and specialist review can answer different questions. The important issue is whether the arrangement catches consequential errors while making useful improvements affordable to evaluate. The time required to reach an acceptable result, and the defects still missed along the way, belong in that assessment.

5. What experience cannot authorize

Even a well-tested method may be inappropriate for a later task. The governing rule may have changed. A qualification omitted from a summary may now be decisive. Two matters can look similar while requiring different inquiries because the relief sought or the stage of the proceeding differs. A method needs to remain answerable to the work in front of it.

Its source can impose another boundary. Imagine deriving a useful examination technique from confidential witness preparation. Removing names would not necessarily remove the sensitive information. The unusual sequence of events, the choice of questions, or the weakness being investigated could still reveal something about the representation. I would treat permission to reuse the lesson as a separate question from whether the lesson works.

This applies within organizations as well as between them. A person’s access to one matter should not silently become permission for another team to inherit information derived from it. Broad reuse needs an authorized basis. In some cases, the most responsible form of learning will remain local to the original work.

There is a tension here. Methods need enough context to remain legally meaningful, while that context may be precisely what cannot travel. Making a lesson more abstract can reduce both its sensitivity and its usefulness. Sometimes it will be possible to develop a separately tested, general procedure. Sometimes the answer should be to retain the experience without exporting a reusable lesson.

Nor should accumulated experience force every problem into a favored method. Consider a system that repeatedly recommends one approach. That approach will generate most of the subsequent feedback, while alternatives remain poorly observed. A record of repeated acceptance could then reflect what the system exposed people to rather than what would have served them best. Maintaining the ability to compare other approaches is part of learning from experience rather than merely accumulating it.

The applicability question also affects what reaches a later task. WikiSkill deliberately supplied active skills directly during its experiments, leaving retrieval and triggering outside the evaluation.3 A useful method stored somewhere is not yet an improvement in a lawyer’s work. We still need to establish when it is selected, whether its conditions apply, and whether it changes the result beneficially. It must also be possible to stop using it when those conditions no longer hold.

An institution worth teaching

I am optimistic about this direction because the value of an expert correction can outlast the document in which it first appears. A system that preserves the reason for a correction, tests its limits, and makes the resulting method available at the appropriate moment could reduce how often lawyers have to repair the same kinds of mistakes.

The more interesting possibility is that this changes the quality of the questions they can pursue. A team spending less time reconstructing the record could spend more time investigating an alternative explanation. A junior lawyer could encounter the reasoning behind a qualification before repeating the error it was written to prevent. A small team could have more capacity to test its strongest argument against a serious objection. Whether these benefits materialize should be tested in the work itself.

The limiting resource may increasingly be the capacity to establish which proposed improvements deserve to survive. A system that generates a thousand new procedures while leaving a partner to investigate each one has moved the burden rather than removed it. Progress will depend on making the evidence for a useful change easier to inspect, while keeping disagreements and failures available for further inquiry. Recursion alone promises no particular rate of improvement.

Herbert’s warning matters because early choices shape what later experience is allowed to reinforce. A system should be able to preserve a valuable correction without turning it into doctrine, and learn from an expert without treating that expert as infallible. The next lawyer should inherit a better starting point, together with the reasons for changing it.


Notes

Footnotes

  1. Frank Herbert, Dune, opening chapter epigraph, attributed within the novel to Princess Irulan’s Manual of Muad’Dib. See the publisher-authorized excerpt. Back to reference 1

  2. Marc-Antoine Allard, Arnaud Teinturier, Victor Xing, and Gautier Viaud, Experiential Reflective Learning for Self-Improving LLM Agents, ICLR 2026 MemAgents Workshop. See §§2–3 and Figure 2, p. 3. Reported average success on the evaluated Gaia2 Search and Execution tasks increased from 48.3% to 56.1%, a difference of 7.8 percentage points. The experiment used a common agent backbone and separate training and test data environments. This is not a legal-work benchmark. Back to reference 2 Back to reference 2-2

  3. Liyan Tang and colleagues, WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution, 2026. See §§3.1–3.2 for the separation of traces, knowledge, and procedures, including preservation of the knowledge layer after rejected skill changes; §4.1 for benchmark domains; and “Limitations,” p. 14, for the exclusion of skill retrieval and triggering from the evaluation. Back to reference 3 Back to reference 3-2 Back to reference 3-3

  4. Xiyuan Yang, Jiaru Zou, and colleagues, Recursive Multi-Agent Systems, 2026, §§2–3. This version studies trainable connections between models’ internal representations, distinct from application-level written procedures or repeated text exchanges. Back to reference 4

  5. Mingda Zhang and colleagues, SkillFlow: Flow-Driven Recursive Skill Evolution for Agentic Orchestration, 2026, §§4.1–4.3 and Appendix R. The trainable supervisor, frozen executor, reward-proportional objective, and per-step diagnostics are features of the research method. Their interpretation depends on the specified task and reward; they do not establish causal attribution of litigation outcomes. Back to reference 5

  6. Philipp D. Siedler, Strategic Persuasion with Trait-Conditioned Multi-Agent Systems for Iterative Legal Argumentation, 2026. See §3.1 for the ten synthetic cases, §3.6 for the shared judge/advocate model backend, and §6.2 for the stated evaluation and external-validity limitations. Back to reference 6

  7. Li Zhang and Kevin D. Ashley, Mitigating Manipulation and Enhancing Persuasion: A Reflective Multi-Agent Approach for Legal Argument Generation, 2025. See §§3–4 for the factor analyst, argument polisher, predefined trade-secret factors, and arguable/mismatched/non-arguable scenarios; §7.1 for limits concerning factor extraction, depth, automated evaluation, and wider legal generalization. The experiment concerns reflection within a bounded argument-generation task, not persistent learning across matters. Back to reference 7

  8. Manuj Kant and colleagues, Towards Robust Legal Reasoning: Harnessing Logical LLMs in Law, 2025. See §§3–5 for vanilla, unguided, and guided approaches to insurance-coverage queries, and §6 for the narrow policy and evaluation scope. Back to reference 8

  9. Anu Pradhan, Alexandra Ortan, Apurv Verma, and Madhavan Seshadri, LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation, 2025. See §§3–4 for human/model agreement, multiple evaluation dimensions, and the study of 117 legal queries. Agreement with a reference evaluator and correctness of legal analysis are separate propositions. Back to reference 9

Back to top

THE INDEX

Find a thread.

Loading the index…