Process audit · July 2026

How an AI-assisted legal appeal was made and verified

From the first chat draft to v13, filed with a national bar association on 29 July 2026, inside a peremptory statutory deadline. All figures extracted deterministically from the complete session transcripts and cross-verified by ten analyst agents plus an adversarial completeness critic. To our knowledge, this is the first published end-to-end audit of how an AI-assisted legal filing was produced.

Read the case study: the problem, the audit findings, unit economics and the product blueprint.

Case record
62
documents · 211 pages; 35% of the documents exist only as scanned images
Elapsed
3 days
intake to filing, deadline-driven
Versions
v6 → v13
6,084 → 9,289 words (+53%)
Human input
≈7,300
words typed, in 96 interventions
Agents
46
production agents · 6 workflows · 6 agent failures recovered
Model output
6.13M
tokens (993M processed, 95% cache)
Compute cost
$1.46–1.77k
USD, API-equivalent; agentic phase

01Three working regimes

The process ran in three regimes with different human roles. That is the structure a product should replicate.

R1

Chat drafting

before intake · 6 versions

v1→v6 produced by the user alone, drafting the traditional way with a commercial chatbot (ChatGPT, Perplexity or similar) alongside. A solid argument architecture that survived to the filed version, but 18 defects identified later, incl. a quotation that doesn't exist in the source.

R2

Near-autonomous pipeline

27 Jul 17:23 → 28 Jul 01:17 UTC · ≈ 8 h

≈8 hours of overnight agentic work on 16 short human messages: 5 workflows, 26 agents, 45 adversarial attacks judged, a 192-row citation matrix. The core analysis (phases −1 to 4) ran ≈2h10 on just 2 messages. Human = gate operator.

R3

Paired review & filing

28 Jul 11:48 → 29 Jul 22:19 UTC · ≈ 2 days

The reviewer reads in Word and pushes back point by point; the AI verifies against sources before applying transactional edits. 82% of all human interventions; v8→v13; manual signature and filing. Human = paired reviewer.

02Session timeline

Active duration and model output per session (UTC). Sessions fec87fe3 and f4937fdb are the same conversation, forked by a client restart at 18:53:34 on 27 July.

SessionRegimeWindow (UTC)Active time Output (tokens)User msgsMilestones
4abaf260R227 Jul 17:23–17:37
14 min
20k
3Setup: folders, rules, mission charter, v6 docx→md
fec87fe3R227 Jul 17:37–18:47
70 min
229k
2Phases −1 to 3 (workflows); dies in the restart
f4937fdbR227 Jul 17:37–20:03
145 min
271k
3Phases 2–4 complete; human-requested pause honored
3b3dbaa6R227 Jul 20:03–20:30
27 min
164k
2Phase 5: technical opinion + 15 corrections, clean context
8cf1c94cR227 Jul 20:33–01:17
284 min
701k
19v7 Phase 6; account-limit outage (118 min); visual page checks
81391354R328 Jul 11:48–13:59
131 min
372k
10v8 Service-of-notice deep dive; the decisive "hand delivery" finding
7ae18c32R328 Jul 14:02–09:43
≈219 min
734k
26v9 Sentence-level analysis of scanned rulings; +1,826 words
85a08cecR329 Jul 10:27–20:30
≈600 min
977k
45v10v11v12 Review marathon; dies in API 500/529 errors
63def8dbR329 Jul 20:29–…
≈135 min
399k
13v13 Final sprint, 9 verifier agents, filed 22:19 UTC

Not shown: 4da11111 (empty transition residue) and ee148592 (the audit session itself). Zero context compactions across the whole project: state always lived in files.

03Document genealogy

Eight versions, none skipped. Three arcs: citation fixes (v7–v8), rhetorical restructuring (v9–v10), and mostly subtractive logical hardening driven by the reviewer's questions (v11–v13).

v6
≤27 Jul · chat
Final chat-phase draft · 6,084 words
v7
27 Jul · night
9 mandatory corrections from the technical opinion
v8
28 Jul · 13:10
"Hand delivery" finding + burden-of-proof fix
v9
28 Jul · 14:51
Synthesis & case map; +1,826 words, the biggest jump
v10
29 Jul · am
Two rulings disambiguated; corrupted docx caught & rebuilt
v11
29 Jul · pm
Burden of proof, after an evaluated second opinion
v12
29 Jul · pm
Filing rule fixed + logged micro-edits
v13 ✓
29 Jul · 21:05
11 anchors in the conclusions · signed and filed · 9,289 words

Versioning rule that emerged in production: new version when the file is open in Word (lock file), already under human review, or the change is structural. Otherwise, in-place edits with a dated change-log entry.

04Human interaction

96 interventions, bimodal

In the pipeline the human operates gates; in review they annotate the document at ≈1 msg every 15 min for 10 hours.

R2 · pipeline
16 · median 14.5 words
R3 · review
79 · 71% of the words
Unattributed
1 · outside both regimes

Leverage ratio: 1 typed word ≈ 840 tokens of model output ≈ 43 words of persisted work product. The AI asked exactly one structured question in the whole project. Every production decision went through free-form prose.

What the human actually did

The user judges more than they instruct: 49% of interventions are judgment on work done.

Judgment: skepticism (20), factual fixes, scope · 47 Instructions · 18 Authorizations · 12 Other · 19

"Is everything really finished? Even the agents that died mid-run with API errors?" This was the question that forced an exhaustive re-verification and, in the chain it opened, the discovery of a fabricated citation less than 24 h before the deadline.

05Orchestration: self-reported vs verified

The system's own end-of-project summary claimed numbers the disk inventory corrects. That is why honest agent/failure telemetry is a core product feature.

MetricSelf-reportedVerifiedNote
Production workflows56+ an exhibit-verification workflow on filing day (9 agents)
Workflow agents263526 on 27 Jul + 9 on 29 Jul
Standalone subagents4211production only
Total agents684659 transcripts − 13 duplicates from a session fork
Failed agents064 in a client restart + 2 on account limits; all recovered

Models: all 35 workflow agents ran on the frontier model (Claude Fable 5); the main thread switched to Claude Opus 5 for the detailed review sessions (at the exact moment of the user's skeptical challenge) and back for the filing sprint. Occasional cross-checks ran on third-party models (GPT-5.6 Terra, Gemini 3.1 Pro), the source of the "second opinion" evaluated for v11; those tokens are outside every count on this page.

06Tokens and compute cost

Output by model

6.13M output tokens; 993M processed in total, of which 95% is cache reads (re-read context).

Fable 5
3.41M
Opus 5
2.72M

Fable ran the pipeline and every agent; Opus ran the detailed review sessions. Cost is dominated by context re-reading, not generation. Workflow and caching architecture are the main cost levers, and we control both. The pipeline is not tied to these models: the expensive tier is used where judgment is needed, and verification work moves to cheaper models as they close the gap.

API-equivalent cost (USD)

List prices; range = cache-write pricing at 5-min vs 1-h TTL. Actual usage ran on subscription.

Fable 5
$1,076–1,351
Opus 5
$383–417
Total
$1,459–1,768

Excludes the chat phase (not audited) and occasional cross-checks on third-party models (GPT-5.6 Terra, Gemini 3.1 Pro). At least an order of magnitude below the professional-fee equivalent of the same verified work, and two in the most expensive jurisdictions.

07Failures and recoveries

The cross-cutting pattern: automatic recoveries worked; the critical catches came from human skepticism or independent verifier agents, never from the producer checking its own work. Twelve incidents in all: two of them, the client restart and the account limits, account for the 6 agent failures counted above.

FailureSeverityCaught byOutcome
Fabricated quotation in the document since the chat phase● criticalVision agent, after a user challengeFixed < 24 h before the deadline
Sampled verification presented as complete● criticalUser's skeptical questionExhaustive sweep relaunched, partitioned
Word file corrupted by an XML-level edit● seriousUser, opening the fileRebuilt; permanent format-validation guards
"Can't read scanned PDFs": capability wrongly denied● seriousUser challengeVision pipeline built; the image-only documents transcribed
Account limits kill 2 auditors; 118-min outage● seriousLimit messages in agent logsManual resume; work redone
Inherited factual errors in the document● seriousPaired review + source checksFixed with dated change log
Factual error in the filing email, written by the orchestrator● seriousIndependent verifier agentFixed before sending
Client restart kills 4 agents and forks the session● mediumWorkflow journalAuto-relaunch in ≈20 s
Longest session dies in API 500/529 errors● mediumImmediateNew session; state rebuilt from files
Official-gazette SPA breaks the web-reading tool● mediumImmediateReal-browser extraction became the standard method
Formatting regressions introduced by the AI's own edits● mediumDocx-level auditsRun-level formatting diffs added
Race condition in the exhibit-verification workflow● mediumVerifier (spurious fail)Freeze artifacts under verification

Method & scope. Quantitative layer: a deterministic script over the complete session transcripts (11 sessions, 59 agent transcript files, 46 unique agents once fork duplicates are removed). No model touched these numbers. Qualitative layer: ten analyst agents plus an adversarial completeness critic; self-reported figures corrected against disk evidence. Four scope caveats: the chat phase was not audited and is reconstructed indirectly; the assisting counsel's role is not recorded; cost shown is API-equivalent (actual usage on subscription); approval-prompt friction was not instrumented. Times UTC (local = UTC+1). The case is identified here by a Satrix internal reference; the proceeding's own number and identifying details are withheld. Every figure is from the real matter.

Case study: findings, unit economics and product blueprint · Full internal audit report and annexes available on request.

Satrix.ai is not a law firm and provides no legal advice, representation or other legal services. This document reports technical work on an AI system and takes no position on the merits of the matter it draws on.