Artificial intelligence tools are speeding up legal work, but they are also introducing a new liability risk: hallucination. In legal contexts, an AI hallucination happens when a tool fabricates citations, invents legal terms, misapplies law, or invents obligations that do not exist in a contract. When an AI-assisted workflow signs off without human review, those hallucinations become malpractice liability.
We believe AI belongs in the workflow, not in the final sign-off. Here is what you need to know about hallucinations, how to assess tool accuracy, and how to build governance that keeps your legal team safe.
Key takeaways
- AI hallucinations in legal work are documented and real; accuracy varies by tool type (specialized legal AI ~70-80%, general chatbots much lower).
- Hallucinations slip through most easily in high-volume, low-complexity work (NDAs, intake triage) where humans skim AI output instead of reading deeply.
- Professional responsibility rules (ABA 2024) make the attorney accountable for AI output, regardless of tool accuracy.
- Guardrails include proportional human review, workflow design, audit trails, vendor transparency, and clear escalation paths.
- The safest approach is tool-assisted human review, not AI-first decision-making.
What is hallucination in legal AI?
Hallucination occurs when a language model generates plausible-sounding text that is factually incorrect, unsupported, or invented. In legal work, hallucinations manifest in several ways:
- Fabricated citations: An AI tool cites a case, statute, or regulation that does not exist or misquotes the law.
- Invented obligations: A contract analysis tool flags a clause that is not in the document.
- Misapplied law: An AI tool states that a jurisdiction’s law prohibits something it actually permits.
- Fictional terms: A legal research tool generates a defined term that is not standard and presents it as accepted usage.
- Missing provisions: A contract review tool misses an entire clause, creating a false sense of completeness.
Hallucinations are especially dangerous in legal work because they are often internally consistent and difficult to spot without domain expertise. A lawyer skimming an AI summary may miss a fabricated citation until it is too late.
Accuracy benchmarks: What the data shows
Specialized legal AI tools perform better than general-purpose chatbots, but none are perfect.
- Specialized legal AI (Lexis+ AI, LexisNexis research): approximately 83% accurate on contract extraction and risk flagging tasks.
- General legal research AI (Westlaw AI, general LLMs fine-tuned for legal): approximately 67-75% accurate.
- General-purpose chatbots (ChatGPT, Gemini without legal training): 40-55% accurate on legal tasks, high hallucination rate.
These percentages vary by task type. Extraction tasks (pulling key terms from a contract) show higher accuracy. Legal reasoning tasks (should we accept this payment term?) show lower accuracy. Statutory interpretation is particularly error-prone across all tools.
A 15-20 percentage point accuracy gap between tool types matters immensely. If your intake team processes 200 NDAs per month using a chatbot-based system, you should expect 18-40 hallucinations per month if humans do not review the output.
Where hallucinations slip through most easily
Hallucinations are most dangerous when three conditions align: high volume, low complexity, and skimmed review.
High-volume, low-complexity work includes:
- NDA triage and risk classification.
- Contract intake and metadata extraction (contract type, parties, signature block location).
- Intake form completion and request routing.
- Routine compliance checklist generation.
- Initial document search and retrieval.
When legal teams use AI to accelerate routine work, they often adopt a “skim, don’t read” review mindset. A lawyer or paralegal glances at an AI summary and approves it without checking the source. That is where hallucinations hide.
High-stakes work (litigation holds, M&A diligence, confidential negotiations) is usually reviewed deeply. Hallucinations still occur, but human reviewers are more likely to catch them because they are reading every word.
Professional responsibility: You are accountable
The ABA issued formal guidance in 2024 (Formal Opinion 512 and related ethics rules) making clear that attorneys are responsible for AI output. Specifically:
- You must understand what the AI tool does and does not do.
- You must verify factual claims the AI tool makes, especially citations.
- You must disclose to clients that you are using AI in their work.
- You must maintain confidentiality and data security when sending work to AI vendors.
- You cannot rely solely on an AI tool’s output without human review proportional to the stakes.
Malpractice insurance may not cover errors caused by your own negligent use of AI. If you outsource a contract review to an AI tool without human verification and miss a liability clause because of a hallucination, the liability is yours, not the vendor’s.
How to build guardrails: A practical framework
We recommend a three-layer approach to minimizing hallucination risk.
Layer 1: Tool selection and vendor vetting
Choose tools that are designed for legal work and backed by transparency about accuracy and data handling.
- Demand a data handling policy: Does the vendor train on your data? How long do they retain it? Can you opt out of training? Insist on a Data Processing Agreement (DPA) that explicitly prohibits training on your work.
- Require SOC 2 or ISO 27001 certification: This assures basic security hygiene.
- Ask for accuracy benchmarks: Reputable legal AI vendors publish accuracy metrics on their website or in white papers. If they do not, that is a red flag.
- Test before you commit: Run a pilot on historical, non-sensitive work to measure accuracy yourself.
Major vendors in contract analysis include Icertis, Ironclad, and Spellbook. General legal research tools include Westlaw AI, Lexis+ AI, and specialized platforms like Harvey AI. None should be treated as plug-and-play; all require governance.
Layer 2: Workflow design
Design the process so that humans review proportional to the stakes.
- High-stakes work: 100% human review, AI-assisted. Use AI for initial extraction or flagging, but a lawyer reads the full document and the analysis.
- Medium-complexity work: Structured review. Use a checklist to ensure the reviewer addresses the AI tool’s key findings. Do not let the reviewer simply approve the AI summary.
- Routine work: Spot-check sampling. If an AI tool processes 100 NDAs, a reviewer spot-checks 10-15% and escalates any hallucinations they find.
- Escalation triggers: Define rules for when an AI flag should be escalated to a senior lawyer. Example: “Any AI flag about indemnification, liability caps, or confidentiality scope goes to the contract lead.”
Layer 3: Audit trails and transparency
Log what the AI tool was asked, what it returned, and what a human did about it.
- Maintain records: If you use an AI tool to review a contract, save the AI’s output and the human reviewer’s decision. This creates an audit trail in case of a malpractice claim.
- Track accuracy: Measure the percentage of AI recommendations that a human reviewer accepts vs. rejects. If acceptance rates are suspiciously high (>95%), your review process may be too lenient.
- Disclose use: If you send client work to an AI tool, tell the client and document their consent.
- Monitor tool updates: If a vendor updates their model, re-test accuracy on your historical data to ensure the change did not introduce new hallucinations.
Common pitfalls to avoid
Over-reliance on a single tool: Do not use one AI tool as your source of truth for all contract analysis. Consider using one tool for extraction and another for risk flagging to reduce correlated errors.
Skipping the vendor vetting step: A tool with 80% accuracy and no data security is worse than a tool with 75% accuracy and strong privacy controls. Do not choose tools based on brand name alone.
Setting unrealistic expectations: Communicating to your team that “AI handles contract review” invites careless review. Frame AI as a productivity multiplier, not a replacement for human judgment.
Ignoring the cost of hallucinations: If a hallucination causes a missed liability or a failed audit, the cost is far greater than the time saved. Build that into your ROI calculation.
FAQ
Can we use ChatGPT or general chatbots for legal work?
You can use general chatbots for general legal research or brainstorming, but not for client-facing work or confidential analysis without explicit guardrails. General chatbots are trained on public data and may use your prompts to improve their models. If you use them, assume the data is not confidential. For any work involving client information, trade secrets, or litigation strategy, use legal-specific tools with DPAs and explicit no-training clauses.
How do we test an AI tool’s accuracy before committing?
Run a pilot on 50-100 historical contracts or documents that you have already reviewed and understand. Feed them to the AI tool and measure accuracy by comparing the AI’s output (risk flags, extracted terms, classifications) against the ground truth from your human review. Calculate precision (how many of the AI’s flags are correct?) and recall (how many of the actual risks did the AI find?). Acceptable benchmarks are 80%+ precision and 85%+ recall for medium-stakes work.
Should we tell clients we are using AI?
Yes, if the AI tool is making decisions on their behalf (even if humans review the output). Include a disclosure in engagement letters or SOWs. Example: “We may use AI-assisted tools to accelerate contract review and analysis. All AI recommendations are reviewed by a qualified attorney before delivery.” Document consent.
What should we do if we discover an AI hallucination in work we have already delivered?
Notify the client immediately, acknowledge the error, explain the corrective action, and review all related work for similar hallucinations. Document everything. This is a professional responsibility issue. The longer you wait to disclose, the worse the liability looks.
How do we measure whether our AI governance is working?
Track: (1) percentage of AI recommendations accepted vs. rejected by human reviewers, (2) number of hallucinations found in QA sampling, (3) time saved per task (compared to no-AI baseline), (4) client and team satisfaction. If acceptance rates are >95%, review is too lenient. If hallucinations are found frequently, the tool or the review process is failing.
Where to start
AI is a powerful tool for legal operations, but it is not a shortcut to accuracy. We recommend starting with lower-stakes, high-volume work (NDA triage, intake routing) where hallucinations are easier to catch and lower-impact. Build your governance muscle on that work before moving to higher-stakes applications.
We partner with legal teams to design AI-enabled workflows that keep confidentiality intact and human judgment in the loop. If you are evaluating AI tools or want to audit your current AI governance, we can help you build a risk-conscious roadmap.
Learn more about how we approach AI in legal operations. Reach out to discuss your specific use case.