The material fact is often the disagreement
A legal-credit investigation may include a trust deed, amendments, resolutions, identity documents, facility papers, valuations and correspondence. The useful outcome is not a smooth summary of each file. It is a case pack that shows which parties, powers, restrictions and dates agree, which conflict, what is missing and what a lawyer must decide.
A summarisation-first design tends to merge repeated statements and choose the most fluent interpretation. That is exactly where it can fail. Two documents may use the same name for different legal capacities. A later amendment may alter one power while leaving another intact. A signature can be visible but not verified. A restriction may apply only after a triggering event.
The architecture should preserve competing propositions until an authorised person or approved precedence rule resolves them. Agreement is evidence. Disagreement is also evidence.
Establish the case boundary
The case service begins with the investigation purpose, subjects, document classes, jurisdiction, reviewer role and decision boundary. It records which downstream decisions are explicitly outside scope. A legal investigation supporting lending should not become an automated lending decision.
| Boundary field | Purpose | Failure prevented |
|---|---|---|
| Investigation question | Defines the propositions the case must assemble | General document chat with unbounded output |
| Parties and capacities | Separates person, trustee, director, guarantor and beneficiary roles | Same name merged across capacities |
| Jurisdiction | Selects approved legal taxonomy and review route | Foreign form interpreted under wrong regime |
| Document set | Establishes admitted evidence and missing classes | User upload treated as complete file |
| Time frame | Distinguishes execution, amendment and current effect | Later document applied retrospectively |
| Reviewer authority | Names who can resolve which question | Operations user accepting legal interpretation |
| Downstream exclusions | Prevents case pack from triggering other decisions | Extraction result changes facility state |
The case envelope accompanies every retrieval and extraction request. If a discovered document concerns another entity or purpose, the system proposes a scope change. It does not add the material silently.
Treat every page as a governed evidence object
Legal documents can contain scans, stamps, handwriting, redactions, tables, schedules and embedded images. Rendering and OCR are transformations. The manifest must retain native file, version, checksum, page image, text layer, extraction service and quality state.
An extracted signature region can establish that a mark appears at a location. It cannot establish identity, authority, intent or legal execution without the required evidence and review. The system should use evidence-role labels such as observed, extracted, verified, asserted, inferred and legally accepted.
| Evidence role | Example | Permitted wording |
|---|---|---|
| Observed | A mark appears in the signature block | “A signature mark is present on page 14” |
| Extracted | OCR reads a party name | “The extraction returned this name with stated quality” |
| Verified | Identifier matches an approved identity source | “The document name matches the verified identity record” |
| Asserted | Correspondence states that approval was granted | “The sender asserted approval” |
| Inferred | Dates suggest an amendment sequence | “The sequence may indicate; review required” |
| Legally accepted | Lawyer records the interpretation | “The authorised reviewer accepted the stated treatment” |
The NIST Generative AI Profile provides a useful risk-management frame for content integrity, privacy, monitoring and human oversight. The W3C PROV-O standard helps record transformations. Neither replaces legal evidence policy.
Normalise without erasing legal form
The proposition model captures subject, capacity, predicate, object, source span, event and valid time, conditions, negation, certainty and review state. It retains original wording beside normalised concepts. A controlled vocabulary improves comparison, but the original clause remains openable.
Negation and condition handling are material. “May not dispose without consent” differs from “may dispose with consent.” A clause that becomes effective on an event should not be represented as continuously active. Tables and schedules may define exceptions to the main body.
The normaliser should abstain when a capacity, cross-reference or conditional scope is unresolved. A model can propose candidate relations, but the case state distinguishes model proposal from admitted proposition.
Normalisation exists to compare evidence, not to make legal language disappear.
Reconcile propositions across documents
The reconciliation engine groups propositions by party, capacity, subject matter and time. It applies approved equivalence and precedence rules only where they are genuinely deterministic. Everything else becomes a conflict or review question.
| Reconciliation relation | Meaning | Machine action |
|---|---|---|
| Equivalent | Same proposition under approved normalisation | Group while retaining sources |
| Corroborating | Independent evidence supports related fact | Show dependency and source independence |
| Superseding | Later valid provision replaces earlier one for stated scope | Preserve both with valid-time relation |
| Narrowing | Later text restricts scope or conditions | Mark affected propositions for review |
| Conflicting | Propositions cannot both apply in stated frame | Block single-answer generation |
| Missing | Required proposition has no admitted support | Add evidence request |
Precedence should not be inferred from document date alone. An amendment may be ineffective, limited or dependent on another approval. The system can present likely sequence and referenced clauses, but legal effect remains a review state.
Use a minimal replayable case state
The case record should not be a conversation transcript. It contains source manifest, propositions, relations, gaps, reviewer decisions, tasks and output versions. Derived summaries are disposable views. If a source or interpretation changes, the state identifies affected propositions and regenerated sections.
Every transition has a machine-readable condition. “EvidenceReady” means the document matrix and extraction-quality gates pass, not that a model believes it has enough context. “Accepted” requires a named reviewer and decision object. “PackIssued” binds the exact evidence and interpretation versions.
The state also preserves unknown outcomes. If a document-management call times out, the service cannot conclude that no document exists. It records attempted source, time, error and safe retry behaviour.
Draft the case pack from structured state
The case pack should lead with the investigation question, parties and capacities, material chronology, agreed propositions, conflicts, missing evidence and reviewer decisions. It can contain a readable narrative, but each material statement points to the case graph.
Protected values include names, capacities, dates, amounts, asset identifiers, clause references and accepted outcomes. The language model cannot alter them. The validator also checks prohibited speech acts such as “legally valid,” “authorised” or “no restriction” unless the statement is directly bound to the appropriate reviewer decision.
Research on structured legal extraction is relevant but should be read narrowly. De Jure studies machine-readable rule extraction and explicit evaluation. From Regulation to Requirements evaluates traceable requirement derivation. These results encourage typed extraction and source linkage; they do not establish legal judgement for a case.
Evaluation must include contradiction and abstention
Field accuracy is necessary but not sufficient. The evaluation set should contain amendments, duplicate names, different capacities, missing schedules, low-quality scans, handwritten changes, conflicting dates, ineffective documents, cross-references and malicious embedded instructions.
| Dimension | Measure | Material failure |
|---|---|---|
| Document integrity | Correct file, page and version | Amended document omitted |
| Extraction | Field and clause accuracy by form and scan quality | Wrong capacity assigned to party |
| Normalisation | Concept mapping with original text retained | Negation removed |
| Reconciliation | Conflict recall and false-merge rate | Two people or capacities merged |
| Temporal logic | Effective, executed and superseded treatment | Later clause applied to earlier event |
| Pack integrity | Claim support and protected-value match | Summary chooses one unresolved proposition |
| Human workflow | Review quality, correction trace and turnaround | Sign-off recorded without material conflict view |
The negative cases matter most. A well-designed system earns credit for refusing to conclude and for presenting the exact unresolved issue. Metrics should separate extraction confidence from legal acceptance. A high-confidence parser does not reduce the authority needed for interpretation.
Production monitoring tracks document-quality mix, extraction corrections, repeated conflict classes, unresolved gaps, reviewer overrides, reopened packs, access denials and source latency. Corrections feed component improvement only through approved, privacy-aware datasets.
Operate the service as a legal evidence workbench
Document operations own intake and source quality. Legal domain owners define proposition vocabulary, review questions and accepted outcomes. The platform team owns secure processing, case state, lineage, evaluation and observability. Information security and privacy govern access and retention. Credit or business teams consume only the approved pack required for their decision.
The value case measures accepted reduction in evidence assembly, duplicate review removed, conflict detection, first-time-complete packs, review rework and downstream defects. It should not use pages summarised as a productivity metric.
Work a case through capacity, authority and time
Consider an investigation into whether an instruction was authorised for a trust-related account. The case contains an account opening form, two trust deeds, an amending deed, a power of attorney, correspondence, certified identity documents and a later operational note. Several people share a surname. One person appears as trustee in one document, beneficiary in another and attorney under a separate instrument. The operational note states a conclusion but does not contain the legal basis.
The service first establishes document identity and relationship. Hashes distinguish copies from altered files. Execution blocks and certification marks are linked to the page image. The amending deed is associated with the instrument it changes rather than treated as another independent source. Party resolution creates separate person and capacity nodes. A person match does not collapse trustee, beneficiary and attorney roles.
The material proposition is not “the person is present in the documents.” It is whether that person, acting in the relevant capacity, held the required authority at the transaction time and whether conditions or joint-action requirements applied. The system displays supporting and conflicting spans side by side. It does not promote the operational note over executed instruments because the note is newer.
The lawyer records interpretation in a typed decision object that cites the evidence, names unresolved assumptions and states effective time. The generated narrative renders that decision and the underlying conflict register. The pack is valuable because it preserves the path to judgement; the prose is only one view of that path.
Resolve entities without collapsing legal capacity
Conventional entity resolution optimises whether records refer to the same real-world person or organisation. Legal investigation needs an additional layer: in what capacity does the entity appear, under which instrument, for which property or account, and during what interval? Identity and capacity must remain separate relations.
| Object | Example attributes | Evidence rule |
|---|---|---|
| Person | Name variants, date of birth, address evidence | Never merge on name similarity alone |
| Organisation | Registered identity, jurisdiction, historic names | Preserve legal-entity version and status |
| Capacity | Trustee, attorney, director, beneficiary | Cite instrument and operative clause |
| Instrument | Type, execution state, effective interval | Retain original page and amendment chain |
| Authority proposition | Action, limits, conditions, joint requirement | Derived only from admitted capacity evidence |
| Transaction or instruction | Time, channel, amount, destination | Link to authority question without presuming validity |
Candidate matches can use names, addresses, identifiers and relationship context, but protected merges require deterministic checks or human confirmation. A merge decision records the compared evidence and can be reversed. Downstream propositions retain the source identities that existed when they were created, allowing controlled recalculation after correction.
Capacity vocabulary also needs jurisdiction and instrument context. A label that looks equivalent across documents may carry different powers. The system should expose source wording and approved taxonomy mappings. It should not normalise away limitations such as “jointly,” “only after incapacity,” or “excluding disposal.” A shared person identifier is not evidence of shared authority.
Model clause precedence and temporal effect explicitly
Documents often disagree because they apply to different periods, subjects or conditions. The reconciliation layer should represent execution date, stated effective date, filing or registration date, knowledge date and supersession relation. These are not interchangeable.
Clause relations can include amends, replaces, survives, is-subject-to, defines and incorporates-by-reference. An amendment may replace one clause while leaving the rest of the instrument in force. A schedule may govern a particular asset. A definition may alter the apparent meaning of a later operative provision. Extraction that ignores these relations creates fluent but legally incoherent summaries.
The system proposes a precedence graph and highlights unresolved cycles or missing referenced materials. Deterministic checks catch impossible dates, duplicate execution blocks and absent schedules. A legal reviewer confirms interpretation. If the case lacks a referenced document, the conclusion records that dependency rather than guessing from common drafting practice.
Every proposition has valid time and knowledge time. A later-discovered amendment may change the legal position at the historic transaction date while also explaining why the institution acted differently based on what it then possessed. The case pack can display both without conflating legal effect, operational knowledge and hindsight.
Protect privilege, purpose and disclosure boundaries
Legal workbenches can join material from customer records, counsel advice, disputes, investigations and ordinary operations. Retrieval must enforce matter, purpose, ethical wall, role and subject before content reaches a model. Search relevance is not an access decision.
The access manifest records which repositories, matters and document classes were queried, which filters applied and which evidence objects were admitted. It need not expose restricted content to unauthorised support staff. Operational traces use identifiers and redacted metadata. Prompts, caches, evaluation corpora and reviewer exports inherit the most restrictive applicable handling rule.
Documents are untrusted input. Hidden text, comments, tracked changes, embedded objects, QR codes and external links can carry instructions or data that should not control the agent. The parser separates content from commands. Tool calls are planned from the case workflow, not from document text. Outbound network access is blocked unless a named research step is authorised.
Disclosure is its own controlled operation. An investigation pack suitable for internal legal review may not be suitable for a customer, regulator or litigation hold response. The destination policy selects fields, redactions, approval and watermarking. Permission to analyse evidence does not imply permission to disclose the resulting synthesis.
Handle corrections, reopening and unknown outcomes
A source can be replaced, a page rescanned, an identity merge reversed or a legal interpretation revised. The platform should never overwrite the historic case silently. It creates a new evidence or decision version, determines affected propositions and marks dependent packs stale until reviewed.
| Change | Automatic response | Required authority |
|---|---|---|
| Better scan of same page | Re-extract and compare material fields | Document operator accepts identity |
| New executed instrument | Add source; invalidate dependent propositions | Legal reviewer determines effect |
| Entity split | Recompute capacity and relationship links | Authorised identity reviewer |
| Interpretation correction | Create superseding decision object | Legal authority for the matter |
| Disclosure already sent | Preserve receipt; open remediation workflow | Disclosure and incident owner |
| Tool timeout during export | Reconcile destination before retry | Workflow controller under idempotency policy |
An ambiguous export must not be repeated blindly. The workflow uses a case and destination idempotency key, then reads the destination or requests operational reconciliation. The same discipline applies when creating tasks or placing legal holds. Unknown is a legitimate state with an owner and service target.
Reopened cases should show the exact reason, affected sections and earlier decision. This avoids asking a reviewer to reread every page. Production measures include correction age, stale-pack exposure, reopen rate, evidence conflicts discovered after sign-off and disclosure incidents.
Review the architecture before trusting the narrative
The design authority should be able to answer:
- How are document copies, amendments, schedules and superseded instruments distinguished?
- Can one person retain several capacities without those powers being merged?
- Which dates drive legal effect, operational knowledge and case cut-off?
- Where are negation, conditions, joint-action rules and exceptions preserved?
- Which gaps force abstention or a legal review question?
- How are privilege, matter scope and destination policy enforced before model access?
- Can an embedded document instruction change tools, scope or disclosure?
- What is invalidated when a source, identity or interpretation is corrected?
- How are ambiguous exports reconciled without duplicate disclosure?
Build an evaluation corpus around material legal failure
Evaluation cases should be designed around consequences, not document neatness. The corpus needs common instruments and difficult combinations: amendments without originals, missing schedules, two people with the same name, one person in several capacities, inconsistent dates, handwritten alterations, partial scans, translated material, conflicting certifications and later operational assertions.
Gold records should include accepted propositions, alternative interpretations, required questions and evidence spans. They should not force one answer where qualified lawyers consider the point genuinely uncertain. The expected behaviour can be to preserve alternatives and escalate. Component scores then distinguish extraction, identity, capacity, temporal relation, conflict detection, evidence support and destination control.
Mutation tests alter a capacity, negation, date, joint-action condition or amendment relation. The case conclusion should change or become unresolved. Access tests place relevant evidence outside the user's authorised matter and confirm that it is neither retrieved nor revealed through citations. Injection tests put apparent tool instructions in comments, scanned text and attachments.
The NIST AI Risk Management Framework helps structure ownership and measurement. The NIST Privacy Framework is useful for data processing and disclosure design. The NIST Secure Software Development Framework supports the protected build and release path. These sources inform the engineering controls; they do not determine legal interpretation.
| Release question | Evidence required |
|---|---|
| Can the service preserve source integrity? | Hash, page and version accuracy across difficult files |
| Can it keep identity and capacity separate? | False-merge and false-split results with material cases |
| Can it recognise unresolved conflict? | Recall of contradictions, gaps and legitimate abstentions |
| Can it resist untrusted document instructions? | Tool, retrieval and disclosure attack tests |
| Can a reviewer correct the state safely? | Invalidation and replay evidence after controlled changes |
| Can an export be governed? | Destination approval, redaction and postcondition tests |
Production sampling should include apparently clean packs, not just escalations. Review reopened cases and downstream corrections because they reveal failures missed at sign-off. Monitor performance by document type, scan quality, jurisdiction and case complexity. A rising abstention rate can indicate source deterioration or a safer response to new document forms; it requires diagnosis rather than a blanket target.
The evaluation standard is whether the service preserves the decision boundary under difficult evidence, not whether its summaries sound legally polished. This makes release decisions understandable to legal, risk, security and engineering owners.
Design the reviewer experience around questions, not documents
A legal workbench should reduce navigation without hiding source form. The main view can present the case question, parties and capacities, timeline, proposition ledger, conflicts, gaps and required decisions. Every item opens to the exact page, image region and extraction state. Reviewers can compare relevant clauses without losing the surrounding instrument.
The interface should distinguish source text, extracted field, normalised concept, inferred relation and accepted legal judgement through consistent visual treatment. Confidence belongs to the extraction or match that produced it. It should not colour a lawyer's accepted decision as if probability and authority were the same thing.
Bulk approval is unsafe for material propositions. Review can be accelerated through grouped evidence and repeated form structures, but the system shows which items share a basis and which carry exceptions. Keyboard operation, readable tables, screen-reader labels and printable evidence views are part of professional usability, especially for long cases.
Corrections are made at the earliest wrong object. If a name was read incorrectly, fix the extraction. If two identities were merged, reverse the merge. If evidence is correct but interpretation differs, record a new legal decision. Editing only the final paragraph leaves the corrupted structured state available for reuse.
The reviewer should also see why a source is absent. “No evidence found” is different from “repository not searched,” “access denied,” “document reference missing” and “query completed with no match.” That distinction prevents false certainty and directs follow-up to the right owner.
Establish an operating and economic model
Capacity planning should use case complexity, not page count alone. A short bundle with competing capacities can require more legal effort than a long standard agreement. Cohorts can be classified by document diversity, entity ambiguity, amendment depth, scan quality, jurisdiction and decision consequence.
| Measure | Better interpretation | Weak proxy |
|---|---|---|
| First-time-complete evidence pack | Required evidence and questions present at review | Pages processed |
| Conflict discovery | Material disagreement surfaced before decision | Extraction volume |
| Reviewer correction by layer | Identifies component quality and rework | Summary acceptance alone |
| Reopened case rate | Tests durability of source and judgement | Pack generation speed |
| Time by complexity cohort | Shows useful workflow improvement | One average handling time |
| Downstream defect | Measures consequence after sign-off | Model confidence |
The product team owns service quality and backlog. Legal operations own case intake and workflow. Domain lawyers own proposition vocabularies and decision authority. Information governance owns retention, purpose and disclosure. Engineering owns secure processing and recovery. Independent assurance chooses its own samples and conclusions.
Value should be claimed from reduced evidence search, fewer duplicate reviews, earlier conflicts, more complete packs and lower downstream correction. Automation of low-value document handling can be substantial, but it is not the same as automation of legal judgement. The commercial case becomes stronger when it prices both saved effort and avoided decision defects.
Rollout should begin with a narrow instrument family and a decision question that has clear ownership. Historical cases are reconstructed without influencing live work. Reviewers compare source integrity, conflict discovery and correction burden. The next phase places the workbench beside the current process while the signed case pack remains unchanged. Production use follows only after access, replay, reopening and disclosure controls are proven.
A service can expand document types without expanding decision authority. It may add better handwriting recognition or a new form parser while legal sign-off remains constant. Conversely, a proposal to reuse an accepted proposition in a downstream credit or servicing decision is an authority change and needs separate evaluation. Capability and authority should never be bundled into one release label.
Operational incidents require case-level impact analysis. If an extraction release mishandled a negation, the team identifies affected pages, propositions, packs and disclosures. It does not simply deploy a fix. Owners decide whether to reopen, notify or remediate each material case. The preserved evidence graph makes this possible.
This controlled expansion gives legal teams a useful workbench before every difficult interpretation problem has been solved. It also keeps the evidence standard intact as volume grows.
Legal document intelligence is strongest when it makes disagreement harder to miss. A useful system gives the lawyer a smaller, better-organised problem with openable evidence. It does not manufacture certainty or transfer legal accountability to the model that wrote the summary.