Legal research that requires a source for every proposition
Fabricated case citations have already led to sanctions for lawyers. In this research system, every proposition must contain an exact passage from the source text or fail validation.
The system searches a saved collection first and builds the answer only from passages it finds. Every proposition must point to an exact passage in the source text.
5 min read
Lawyers have already been sanctioned for briefs that cited cases generated by a chatbot and never existed. A fabricated citation becomes part of a court filing and creates direct professional risk for the attorney.
Meanwhile the need is real. Case-law research eats hours per matter, because the lawyer's question is argumentative and the search tools are keyword-based. A lawyer looks for propositions supporting a line, not documents containing a phrase.
An order that binds every proposition to a source
In the reading room for arguments we built, the system first searches a saved collection of rulings and then assembles propositions from the passages it finds.
The result format enforces the rule. Each CitedProposition must include an exact passage and a pointer to its location in the source document. A proposition without a retrieved passage fails validation before it can appear in the answer.
Verification is one click: the quote expands to the passage in the ruling, with a canonical link to the public record. The saved collection includes source addresses and checksums, so every answer can be traced back to its source.
Search quality with an explicit diagnostic
A law firm often asks about local processing first, so the entire search runs on site, with no external API. It combines BM25 keyword search, multilingual semantic matching, and a local step that reorders the results.
Local search still needs critical evaluation. On seven questions where editorial review found at least one relevant passage, BM25 keyword search scored 0.894 for the ordering of the first ten results (nDCG@10), while the combined method scored 0.870. The combined method did not improve this measure. A separate check across ten questions scored 0.637 for BM25 and 0.713 for the combined method, but it measured membership in the source query rather than legal relevance to an argument.
What those retrieval scores establish
nDCG@10 evaluates the order of retrieved passages against a human relevance judgment. It helps answer whether useful material appears near the top of the list. It cannot establish that a proposition is a sound interpretation of law, that a ruling is still current, or that the passage fits the facts of a particular matter.
A production evaluation should separate three layers:
| Layer | Control question | Owner |
|---|---|---|
| Retrieval | Did the system find the relevant passages in the approved collection? | a lawyer preparing the evaluation set |
| Citation integrity | Does the displayed quote occur exactly in the identified document? | system validation plus sample review |
| Legal usefulness | Does the passage support the argument and remain current? | the lawyer conducting the research |
Combining all three into one score makes failures hard to diagnose. A weak answer may come from a missing document, poor result ordering, or a research question that needs reframing. Each calls for a different correction.
Preparing a firm's source collection
Most of the work before a pilot concerns source governance. Each document needs a stable identifier, source address, retrieval time, checksum, and version. The team also needs rules for which collections a group may search and how a superseded or withdrawn source is handled.
A useful preparation list covers:
- the courts, periods, and matter types included in scope,
- chunking rules that preserve enough context around each passage,
- questions taken from real research tasks,
- questions that intentionally have no answer in the collection,
- passages marked relevant plus deceptively similar negatives,
- the update process and the tests that run after a collection change.
No-answer questions are essential. They show whether the system can stop with an explicit gap instead of assembling a plausible proposition from adjacent material.
An explicit end to research
The system has one more boundary, as hard as the citations: a request for legal advice ends in a refusal recorded in the audit. On 40 frozen public SAOS excerpts, all three required refusals fired and all 14 citation spans matched the source text, with zero external API calls.
The test covers 40 excerpts and 10 questions, including three with no answer in the corpus. It verifies citation spans and refusal behavior for this frozen material. A lawyer still owns interpretation, source currency, and the judgment about whether a passage supports the matter at hand.
Acceptance checklist for a legal team
Before the system supports live matters, verify:
- Every proposition opens the exact quote and full source document.
- The user can see the searched collection and its update status.
- A missing relevant passage produces an explicit gap.
- Ruling text, metadata, and model commentary are visually distinct.
- A previous analysis can be reconstructed against the source version it used.
- Requests for advice, outcome prediction, or filing-ready text follow the defined human-review path.
- Collection updates rerun retrieval, citation-integrity, and refusal checks.
The practical pilot measure is usefulness to the lawyer: time to reach a relevant passage, the share of propositions rejected during review, and the reason for each rejection. Exact citation integrity is a prerequisite, but it does not by itself prove the strength of a legal argument.
Details and measurements are on the case page. If research consumes hours and your team needs checkable sources, a free process scan can map a source-bound workflow for your corpus.
Free process scan
Start with a free process scan.
- A 30-minute call with the engineer who would lead the work.
- A review of the processes that cost you the most time and money.
- A written summary of what to automate first and the likely cost range.
The scan is free and creates no obligation. If automation is unlikely to pay off, the written recommendation will say so.
€0
30 minutes · written takeaway within 2 business days
Times are shown in your own time zone. We work with clients across time zones.
Describe the process in the form