Ten research-grounded predictions for legal AI through the end of 2026 — from the first disbarment for hallucinated citations to the collapse of point-solution vendors to the pricing collision between AI-enabled firms and their clients.
A new working paper ran generative AI document review and a managed active-learning workflow head to head on the same 45,004-document corpus. Same review protocol, same reference labels, same scorecard. The GenAI system won on recall, and the paired test backs it up. The rest of the scorecard — precision, the 62-to-1 effort gap, the population extrapolation — needs more qualification than the headline suggests.
The APEX benchmark — built by Mercor, with tasks authored by BigLaw-experienced lawyers and advised by Cass Sunstein — is the most rigorous test of whether AI can perform real legal work. The answer is more specific than vendors or skeptics suggest.
An Instagram account-takeover wave exploited Meta's AI support bot at the password-reset gate. The lesson for law firms: authentication and ethical walls exist to refuse persuasion — exactly what agents are built to do well.
Excel custom number formats let a cell store one value and display another. Every extraction library reads the stored value. Every LLM platform I tested shifted from 'do not pursue' to qualified interest on the same file.
Every headline called Kirkland's $500M commitment an AI bet. The signals in the announcement — no named model, full exclusivity, value-based pricing — point to something different: an infrastructure play that happens to run AI.
AI processing costs ~3% of an AI-enhanced eDiscovery workflow. The real savings come from restructuring leverage — shifting volume QC from $750/hr associates to $50/hr contract attorneys. Here's the math.
A walkthrough of building a Medicare fraud backtest overnight in Claude Code — from a plain-English spec to 289 matched providers across 41 states, a fraud-similarity model with AUC 0.79, and a manual public-record check of high-scoring peers. Including the three times the pipeline failed, the data duplication bug, and the engineering decisions that shaped the final design.
The attack surface isn't AI — it's the documents AI processes. Prompt injection in discovery, adversarial inputs delivered through Rule 34 productions, and the cybersecurity gaps firms create by piping untrusted content through LLM pipelines.
The previous post described a Medicare fraud backtest nobody had built. Here are the results. 289 excluded providers across 41 states, matched to pre-exclusion billing data, compared against 3.39 million peers. Thirteen of fifteen features showed statistically significant differences — and the same behavioral fingerprint shows up in never-excluded providers who have independent enforcement histories.