How to watch a legal AI demo: the questions that cut through
Every legal AI demo is designed to impress you. Almost none are designed to inform you. A practical guide to the questions that separate a genuinely useful tool from beautifully rehearsed theatre.
I have sat on both sides of the legal AI demo. I’ve given them, and I’ve watched a lot of them, and here is the uncomfortable truth from someone who has done the selling: a demo is a performance, not an inspection. Every document you see has been chosen. Every query has been rehearsed. Every path through the product has been walked a hundred times before you arrived. None of this is dishonest - it’s just what a demo is - but it means the thing you’re watching is designed to impress you, not to inform you.
The good news is that you can turn a performance into an inspection with a handful of questions. The vendors with real products answer them happily. The ones without real products visibly wobble. Here is the list I’d bring.
Ask to use your own documents
This is the single highest-value move available to you, and it’s remarkable how rarely buyers make it. The demo corpus has been curated - clean documents, well-structured, chosen because the tool performs well on them. Your documents are not like that. Your documents are scanned at an angle, half-defined, inconsistently named, and full of the specific mess your practice actually produces.
So before the meeting, send over a handful of your own real (appropriately anonymised) documents and ask for the demo to run on those. A vendor with a robust product will say yes, because they’ve seen messy documents before. A vendor whose product only works on the happy path will find reasons why that’s difficult this time - data protection, configuration, “our sandbox environment”. Sometimes those reasons are genuine. But watch how hard they try to solve them, because that tells you which kind of vendor you’re dealing with.
Ask to type the query yourself
The rehearsed query is rehearsed because it works. The interesting question is what happens one inch off the rehearsed path. So ask for the keyboard. Phrase the question the way an actual associate would phrase it at 6pm - vaguely, with a typo, assuming context the tool doesn’t have. Ask a follow-up that depends on the previous answer. Ask something at the edge of the tool’s scope and see whether it says “I can’t do that” or produces something confident and wrong.
You’re not trying to trip anyone up for sport. You’re testing the property that actually governs adoption: does this thing survive contact with a normal user on a normal day? A tool that only works when driven by the person who built it is not a tool yet. It’s a prototype with a sales team.
Ask what happens when it’s wrong
Every legal AI tool is sometimes wrong. This is not a scandal - it’s a property of the technology, and everyone in the room knows it. What matters is whether the product has been designed around that fact, and you can find out with three questions. How do I know when to check an answer? Show me what an uncertain answer looks like next to a confident one. When it cites a source, what exactly happens when I click it?
Strong products have rich answers here - confidence signals, linked passages, visible reasoning, a workflow that assumes verification. Weak products change the subject to their accuracy percentage. If the answer to “what happens when it’s wrong” is “it’s 94% accurate”, you’ve learned that error handling was not part of the design, and the remaining 6% is going to become your associates’ problem at the worst possible moments.
Ask about the week after go-live
The demo shows you minute one of using the product. The value lives in month six, and the road between those points is where most legal AI actually dies. So drag the conversation there. Who at the firm has to do work to keep this tool useful - and how much? What does onboarding a new practice group actually involve? What did your last three implementations look like, and can I speak to one of those firms without a chaperone?
That last one matters most. A reference call the vendor sits in on is another performance. A reference call without the vendor is where you hear the sentence that tells you everything, which is usually some variant of “the tool is good, but…” - and the second half of that sentence is the real product review.
Ask what the tool won’t do
This question sounds soft and is actually the sharpest one on the list. A vendor who knows their product deeply will answer instantly, because knowing a tool’s edges is what it means to know a tool. “It’s not good at manuscript amendments.” “Don’t use it for advice, it’s a review tool.” “It struggles with non-English documents.” Immediate, specific answers are the sound of a team that has watched real users hit real limits.
The answer to worry about is the one where every limitation dissolves into a roadmap. “That’s coming in Q3” is sometimes true. But a product with no admitted weaknesses is a product whose weaknesses you’ll be discovering yourself, live, on client work.
The demo is the ceiling
One final calibration to carry into every one of these meetings: whatever you’re shown is the best the product currently looks. The demo is the ceiling, not the floor. Real usage - your documents, your people, your deadlines, your mess - starts somewhere below what you saw, and the questions above are how you find out how far below.
None of this requires technical knowledge, which is the point. You don’t need to understand the model architecture to evaluate a legal AI tool, any more than you need to understand engine design to test-drive a car. You need to take it off the route the salesperson planned, onto your own roads, and pay attention to the noises. Lawyers are professionally trained to cross-examine confident witnesses. A demo is just a confident witness. Treat it like one.
Written by Dom Conte
Legal-tech founder, builder and speaker. More about me →