Start a conversation Contact
← Research

Faster junior drafting needs an apprenticeship control

AI can improve a junior professional's draft before it improves their judgement. Capture the added capacity, but keep a measured part of the workflow where the person must find and repair the defects themselves.

A managing partner sees a junior patent lawyer produce a stronger first draft in less time with an AI assistant. The immediate opportunity is clear. More work can move through the practice without adding the same amount of drafting effort, and partners can spend less time correcting avoidable defects.

The firm still sells judgement. A junior who can produce a good document only while the assistant is present has helped today’s capacity. They have not yet become the reviewer who can find the strategic error in somebody else’s work, explain it to a client and decide what must change.

Faster junior drafting needs an apprenticeship control. Use AI where it improves the work, then preserve a measured set of unassisted review tasks and reasoned partner feedback. The commercial aim is to gain capacity now without weakening the supply of judgement the practice will need later.

Assisted quality and acquired judgement are different results

A September 2026 field experiment makes the distinction unusually visible. The researchers randomly assigned 133 practising patent lawyers at eleven US intellectual-property firms to receive a custom AI drafting assistant or remain in a control group. Blinded patent lawyers at another firm scored benchmark work for enforceability, accuracy, strategic ambiguity, completeness and clarity. Ninety-one participants completed the full three-month protocol. Autor and colleagues, 2026

With the assistant, lawyers produced better patent drafts. The treatment group scored 0.34 standard deviations higher after ten days and 0.38 higher after ninety days. The first task took about ten minutes less against a control-group average of 112 minutes. The estimated time saving at ninety days was similar in minutes but statistically uncertain. Junior lawyers captured most of the measured speed gain and larger improvements in assisted drafting quality.

The final exercise removed the assistant. Participants had to redline a flawed patent application, a task designed to expose the judgement behind the edit. Treated senior lawyers outperformed senior controls. Treated juniors showed no average improvement over junior controls, and their results spread out. There were fewer middling scores, but more poor scores as well as more good ones.

That result does not establish that AI damaged junior learning. The experiment had no baseline test of unassisted skill, lasted three months and could not observe every participant’s tool use. It does establish something a practice can act on: better assisted output is not evidence that professional judgement has developed. The two outcomes need separate tests.

How an AI drafting workflow can create capacity without concealing the skill question Fig. 01
If the pilot measures output only If the pilot measures work and capability
Give every draft to the assistant and count time saved Name the drafting steps AI may assist and the review steps the person must still perform
Treat partner corrections as ordinary rework Code corrections by defect type and record whether the junior can explain the repair
Compare the final document with the old average Blind-score assisted drafts and separate unassisted review tasks
Expand when usage and output quality rise Expand when capacity improves and unaided judgement holds or develops

Put the assistant at the first-draft boundary

Start where the work is bounded enough to evaluate. In a patent practice, that could mean producing dependent claims or a detailed description from an approved invention disclosure and a named set of source material. In another professional service, it could be the first version of a standard analysis or client document. The useful boundary has known inputs, a review step and defects that experienced practitioners can classify.

Before AI, a junior produces the first draft, a senior reviews it and the junior repairs the work. After AI, the junior prepares the inputs, directs the first draft, verifies every substantive statement and explains proposed changes to the reviewer. The senior still decides whether the work is fit to leave the firm. The difference is where the blank-page effort goes and how quickly the draft reaches a reviewable state.

Do not automate the feedback loop away. If the assistant silently repairs each error, the final document can improve while the junior receives no account of what was wrong. Require the reviewer to identify the defect and the junior to record the reason for the repair. That record becomes both coaching material and an evaluation set for later versions of the system.

An earlier randomised experiment with law students found the same reason to separate speed from quality. GPT-4 access produced large and consistent time savings across realistic legal tasks, while improvements in legal-analysis quality were slight and uneven. The participants were students rather than practising lawyers, and the study measured immediate work rather than skill after sustained use. It supports the measurement distinction without resolving the apprenticeship question. Choi, Monahan and Schwarcz, 2024

Make the capacity gain appear in an operating measure

Minutes saved on a benchmark are a mechanism, not a commercial outcome. The practice owner needs to decide where the released effort will go. It might support more matters with the same team, shorten the interval to a client-ready draft, reduce partner correction or move senior attention towards the few decisions that carry the most consequence. Name one before the pilot starts.

Measure matters reaching an agreed review stage per lawyer, elapsed time to that stage and senior review minutes per accepted draft. Read those measures beside substantive defect rates, revision rounds, missed deadlines, client-requested corrections and any downstream event that reveals weak work. A faster first draft followed by more partner repair has moved effort rather than released it.

Then measure capability separately. Give participants periodic, representative work without the assistant and score it blind using the same rubric. Include a review task, because finding and repairing a defect is different from generating plausible text. Segment the results by experience and prior performance. The average can hide the people whose assisted output rose while their unaided judgement did not.

A professional-services pilot that can support a capacity decision Fig. 02
  1. 01 Bound Choose one repeatable drafting step with approved inputs and an accountable reviewer.
  2. 02 Baseline Record elapsed time, review effort, defect types and accepted output before assistance.
  3. 03 Assist Let the junior direct and verify the draft while the senior keeps the release decision.
  4. 04 Test Blind-score assisted output and periodic unassisted review work as separate outcomes.
  5. 05 Decide Expand, narrow or redesign according to capacity, quality and skill by cohort.

The apprenticeship control is the fourth stage. It is not a ban on AI or a memory test detached from real work. It is a small, recurring sample of the judgement the firm expects a practitioner to exercise when the tool is wrong, unavailable or inappropriate. If that capability matters to the service, the organisation needs evidence that it still exists.

What the Inference Institute can help decide

The useful engagement defines the workflow boundary, the quality rubric and the comparison before a platform is selected. It maps where source material enters, where AI can draft, which assertions require verification, what a senior must decide and which unaided tasks represent durable professional judgement.

The result is a pilot specification a managing partner can use: one class of work, eligible users, permitted inputs, review rules, capacity measures, quality measures and a stopping condition. It also makes the operating choice explicit. The practice can use the released effort for additional work, faster service or deeper review, then test whether that outcome actually occurred.

What this does not tell you

The patent study is a working paper published in September 2026. Its sample was small, its participating firms already performed work for Google, several authors were Google employees and Google funded the direct experimental costs. The tasks were benchmarks, while the on-the-job measures covered fewer people and were too imprecise to establish a production effect. The results do not forecast another firm’s capacity, quality or skill development.

The study also does not identify the best teaching method. An unassisted review sample can reveal a gap. Closing it may require worked examples, supervised reasoning, different task allocation or more deliberate feedback. Those choices belong to the practice and the profession.

The managing partner’s next decision is narrower than whether lawyers should use AI. Choose one drafting boundary where better first work could release real capacity. Measure that gain, and keep enough unaided review to see whether the people producing faster drafts are also becoming the people the firm will trust to judge them.

Filed under · Method · Professional services · Legal AI · Skills Inference Institute · 10 Sept 2026

Bring us the question

Reading this because it is on your desk right now?

That is the conversation we are best at. Thirty minutes, a written summary, no obligation.