Home · Writing · Deployment

AML Alert Triage with a Hard SAR Boundary: Evidence Assembly Without Delegating Suspicion

Decision-Grade Agentic Systems

TLDR

  1. A controlled architecture for using agents to assemble, test and present AML alert evidence while reserving suspicious-activity reporting, confidentiality and consequential relationship decisions to authorised people and policy.
  2. Transaction-monitoring systems produce alerts because a rule or model found activity that deserves examination. An alert is not a finding of money laundering.
  3. Harbour Trade Services is a fictional business customer. Three alerts open during the same review period.
  4. FinCEN's October 2025 SAR FAQs clarify matters including structured filings, continuing activity and decisions not to file.
  5. An AML case draws on transaction records, customer due diligence, account relationships, device and channel data, prior authorised cases, entity registries, sanctions or PEP screening, and approved external information.
Figure 1Rules and approved detection models to monitoring and case retentionCausal and control schematic
Rules and approved detection models to monitoring and case retention10 declared states connected by 10 authored relations. The figure supports the section The composite: three alerts that describe one pattern. L0L1L2L3L4
No
Yes
01
Rules and approved detection models
02
Alerts
03
Deduplicate and link
04
Bounded case hypothesis
05
Authorised evidence assembly
06
Human investigation
07
Applicable reporting test met?
08
Document authorised no-file decision
09
Authorised report preparation and filing
10
Monitoring and case retention
Reading. The authored topology makes 10 declared relations across 10 states inspectable. Read it as the control structure for “The composite: three alerts that describe one pattern”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
On this page

An alert, an investigation and a report are different objects

Transaction-monitoring systems produce alerts because a rule or model found activity that deserves examination. An alert is not a finding of money laundering. It is not a decision that activity is suspicious under a reporting rule. It is not a suspicious activity report. Collapsing those objects is the first design mistake in an “agentic AML” programme.

The FATF Recommendations, amended in June 2026, provide a global standard that jurisdictions implement through their own law. Recommendation 20 concerns reporting suspicious transactions, while Recommendation 21 addresses tipping-off and confidentiality. The legal test, reporting channel, deadlines, protected information and accountable role depend on the regime. A model cannot derive one global filing rule from the word “SAR.”

The architecture in this article makes a deliberate governance choice: an agent may triage alerts and assemble evidence, but an authorised human decides whether the institution makes a regulatory report. That boundary may be stricter than the minimum technology rule in some places. It is chosen because suspicion is a consequential legal judgement, because reporting data is highly restricted, and because the decision must remain attributable to a person operating under approved policy.

The agent can reduce search and clerical burden. It cannot acquire the reporting authority of the institution.

The customer, transactions, alerts, organisations and decision path below are fictional. They form a worked control scenario, not a disclosed bank case, a claim about filing practice or evidence that the design has achieved a particular operational result.

The composite: three alerts that describe one pattern

Harbour Trade Services is a fictional business customer. Three alerts open during the same review period. One concerns transfers to newly observed counterparties. A second concerns rapid movement through a recently added account. A third concerns transaction descriptions that do not resemble the customer's historical pattern. The alerts arrive from approved rules and a conventional anomaly model; their thresholds and scores are not reproduced here because no universal setting would be defensible.

The legacy workflow places each alert in a separate queue. Three investigators pull overlapping account history, KYC records and counterparty data. The first closes for insufficient evidence. The second remains open while an ownership record is requested. The third escalates because it contains more transactions. None sees that the counterparties share a registered director and that an earlier customer record may be stale.

The proposed agent does not decide that the activity is suspicious. It deduplicates alerts, creates a case hypothesis, retrieves authorised evidence, resolves candidate entities, builds a timeline and identifies conflicts. The hypothesis says: “The observed transaction pattern may be inconsistent with the recorded purpose and may involve related counterparties.” That statement is testable and does not assert a crime.

Start with the applicable reporting regime

The policy service must select the legal and organisational regime before the agent sees filing-related instructions. A United States depository institution can look to the FFIEC suspicious activity reporting manual, relevant regulation and current FinCEN material. The FFIEC manual says the bank should determine whether to file based on the customer information available and that investigators should document their conclusions and any recommendation.

FinCEN's October 2025 SAR FAQs clarify matters including structured filings, continuing activity and decisions not to file. They should be read with applicable rules, not converted into a generic model prompt. In the United Kingdom, the National Crime Agency SAR guidance describes a different reporting environment and explicitly says the NCA cannot advise an individual organisation whether it should submit a SAR.

Source type What it can establish What it cannot establish
Applicable law and regulation Reporting duty, protected information, recipient and legal test Whether the facts of one case satisfy the test without authorised judgement
Regulator or FIU guidance Expected process, report quality and supervisory interpretation A universal rule outside its jurisdiction and scope
Institutional policy Roles, escalation, evidence standard and control design Permission to contradict governing law
Detection model or rule Why an alert opened and how the signal was produced That activity is suspicious or criminal
Agent synthesis Which authorised facts align, conflict or remain unknown Legal filing authority or proof of underlying crime
Industry guidance Useful practice and effectiveness questions Binding legal or supervisory obligation

The decision record must identify the selected regime and policy version. This avoids a dangerous failure in multinational platforms: an agent using one country's language, threshold or confidentiality treatment for another entity. Jurisdiction is an input to policy enforcement, not a phrase left for the model to infer.

Evidence assembly should preserve facts, claims and gaps

An AML case draws on transaction records, customer due diligence, account relationships, device and channel data, prior authorised cases, entity registries, sanctions or PEP screening, and approved external information. These sources do not have equal authority. A transaction ledger can establish that a transfer posted. A customer declaration can establish what the customer said. A media article can raise a question. None alone proves the purpose of the activity.

The evidence builder should create propositions rather than a fluent story. A proposition has a subject, predicate, object or value, event time, source, extraction method and evidential status. “Account A transferred amount X to Account B at time T” can be linked to the ledger event. “B is controlled by Person C” remains a candidate relationship until the authorised corporate evidence supports it. “The transaction lacks an apparent lawful purpose” is not a raw fact; it is a conclusion that an authorised investigator may reach after examining the available context.

Figure 2Linked alert case to investigator viewCausal and control schematic
Linked alert case to investigator view10 declared states connected by 13 authored relations. The figure supports the section Evidence assembly should preserve facts, claims and gaps. L0L1L2L3L4 01
Linked alert case
02
Propositions needed for investigation
03
Ledger and payment events
04
KYC and ownership records
05
Approved entity and external sources
06
Evidence graph
07
Confirmed facts
08
Conflicts
09
Unknowns
10
Investigator view
Reading. The authored topology makes 13 declared relations across 10 states inspectable. Read it as the control structure for “Evidence assembly should preserve facts, claims and gaps”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

The agent should state why a source matters without inventing motive. It can say a new counterparty is linked by a shared registered director. It cannot say the director created the entity to launder proceeds. It can show that activity differs from the recorded customer purpose. It cannot equate difference with suspicion.

The evidence packet must remain useful even when the agent's proposed interpretation is rejected. A reviewer should be able to inspect the ledger event, registry record and KYC version independently. Every material statement in the case summary should point to that packet.

Define an alert-to-case contract before adding an agent

An alert is normally a product of a specific detection version applied to a defined population over a stated period. That lineage is easy to lose when an orchestration layer converts alerts into prose. The case intake contract should retain the detection identifier, version, execution time, customer and account keys, population scope, observed conditions, source-event references and any threshold or score used for routing. It should also record whether the alert supersedes, reopens or merely resembles an earlier alert.

The contract must distinguish what the detector observed from why a control owner configured it. A rule may observe multiple transfers below a configured amount within a window. Its documented control rationale may be to identify a structuring pattern for review. The event record should not translate that into “the customer structured transactions.” One is reproducible arithmetic; the other is an investigative and potentially legal conclusion.

Object Minimum fields Authority it carries Authority it does not carry
Detection execution Population, input-data version, rule/model version, run time, health status Shows which control ran on which data Proof that data was complete or the subject acted suspiciously
Alert Observed conditions, event references, score or threshold, routing reason Establishes why review was requested A case conclusion or filing recommendation
Link proposal Alert IDs, shared identifiers, time relation, link rule/version Explains why alerts may belong together Permission to merge customers or discard source alerts
Investigation case Hypothesis, scope, owner, jurisdiction, policy version, state Defines authorised work and accountability A regulatory report or customer-treatment decision
Evidence proposition Claim, source, event/valid time, extraction and review status Supports or challenges a case question More certainty than the source provides
Reporting record Authorised decision, final content, approver, submission receipt Records the protected reporting action Automatic authority over payment or relationship action

Case creation should be transactional. The service first validates that every referenced alert exists and that the case identity is consistent. It then creates the case and records links without mutating the original alerts. If later evidence shows that two alerts concern different people, the investigator can split the case. The split retains the original link proposal and the reason for correction. This makes entity-resolution error visible rather than rewriting history.

Idempotency matters because detection systems retry. An orchestration failure should not open the same investigation twice or send the same evidence request repeatedly. Each event and tool call carries a stable idempotency key. Duplicate delivery produces the existing receipt. Genuine new activity creates a new event that can enrich or reopen a case under policy.

The intake service should reject partial detection runs. If the underlying ledger feed was incomplete, an apparently clean counterparty set is misleading. Health status travels with the alert, and the case interface exposes limitations. A control owner decides whether to rerun, continue under an exception, or pause the affected population. The agent cannot convert a pipeline warning into confidence.

Case scope is another first-class field. It states the subjects, accounts, period, products and hypotheses that the investigator is authorised to examine. Tool calls are checked against it. If the investigation uncovers a related account outside scope, the agent proposes an expansion with the evidence that justifies it. An authorised role approves or rejects the expansion. This prevents graph exploration from becoming unrestricted surveillance.

The case state should also show unresolved dependencies. Waiting for a registry record, customer due-diligence update or another investigation is not the same as analyst inactivity. Dependencies have owners, due dates and fallback policy. A case does not close merely because the agent exhausted its tool budget.

The contract makes triage inspectable before it makes triage faster. It preserves the detector's exact contribution, prevents conclusions from entering through labels, and gives every later narrative a stable set of objects to cite.

Preserve event time, knowledge time and decision time

AML investigation is temporal. Transactions occur at event time. Feeds and external records arrive at knowledge time. The institution takes decisions at decision time. Mixing those clocks can create a persuasive but false chronology. An ownership filing discovered after an alert may describe a change that occurred before it; a device association may have been created after the relevant payments; a KYC record may be the current version today but not the version used when activity was assessed.

For each proposition, retain both the time it describes and when the institution could first use it. Evidence assembled for a historical decision must be cut off at that decision's knowledge time unless the review explicitly asks what later information changes the assessment. That protects backtesting and accountability. It also prevents an agent from criticising an earlier reviewer with evidence that did not yet exist.

Figure 3Title harbour composite: three clocks to 2026-04-02 : authorised reporting decision is recordedCausal and control schematic
Title harbour composite: three clocks to 2026-04-02 : authorised reporting decision is recorded9 declared elements supporting the section Preserve event time, knowledge time and decision time. L0 01
title Harbour composite: three clocks
02
2026-03-04 : First payment event occurs
03
2026-03-06 : Detection run creates Alert One
04
2026-03-18 : Registry change becomes effective
05
2026-03-25 : Second and third payment patterns occur
06
2026-03-26 : Alerts Two and Three open
07
2026-03-28 : Registry amendment becomes available to the bank
08
2026-03-29 : Linked investigation begins
09
2026-04-02 : Authorised reporting decision is recorded
Reading. The figure locates 9 declared elements used by “Preserve event time, knowledge time and decision time”. It is schematic, not measured. Schematic derived from the paper's authored topology; no measured quantities.

A timeline generator should not order events by document retrieval time alone. It normalises time zones, distinguishes posted from authorised payment time, retains value dates where relevant and marks uncertain or interval dates. If a source says “in March”, the system does not invent the first day of March. It represents the interval and explains the uncertainty.

Reversals, returns and corrections need explicit relations to the original event. Counting an original transfer and its technical reversal as two outward payments corrupts totals and can exaggerate rapid movement. Likewise, an amended registry record should supersede a field for the relevant period without erasing what the earlier source showed. The case packet presents both the event chain and the net analytical treatment.

The evidence service should answer four separate questions: what happened in the payment system; what the institution knew when each control ran; what later evidence now changes the understanding; and which evidence supported the final decision. Those answers often overlap but are not identical.

This distinction is especially important for continuing or recurrent activity. Current official guidance, including the FinCEN SAR FAQ page, must be interpreted for the applicable institution and regime. The system can identify activity since a defined prior decision and assemble it. It should not infer the reporting treatment or timing from a generic window copied across jurisdictions.

Late evidence triggers impact analysis. If a corrected ledger event changes a material total, or entity resolution shows that a counterparty was the wrong person, the case service identifies affected propositions, narratives and decisions. An authorised reporting role then applies the applicable correction, amendment or supplemental process. The platform never silently edits the record that was submitted.

Temporal integrity is a control against hindsight disguised as intelligence. A trustworthy case shows what was true, what was known and what was decided on each relevant date.

Make provenance survive every transformation

Evidence rarely moves directly from source to investigator. Payment events are normalised; documents are scanned; tables are extracted; names are transliterated; entities are resolved; graphs are built; text is summarised. Each transformation can introduce an error. A citation to the original document is insufficient if the decisive assertion was produced by a chain whose intermediate steps are invisible.

The W3C PROV-O standard supplies a general vocabulary for entities, activities and agents in provenance. It does not determine AML evidence sufficiency. It does allow the architecture to state that a case proposition was generated from specified source entities through identified transformations, with software and human actors recorded.

Figure 4Immutable ledger events to versions, timestamps and reviewersCausal and control schematic
Immutable ledger events to versions, timestamps and reviewers9 declared states connected by 11 authored relations. The figure supports the section Make provenance survive every transformation. L0L1L2 01
Immutable ledger events
02
Currency and event normalisation
03
Registry document
04
OCR and table extraction
05
Case propositions
06
Entity-resolution activity
07
Typed relationship graph
08
Cited investigator view
09
Versions, timestamps and reviewers
Reading. The authored topology makes 11 declared relations across 9 states inspectable. Read it as the control structure for “Make provenance survive every transformation”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Every material proposition carries a derivation path. For a transaction fact, that path can include the raw ledger identifier, normalisation rule and aggregation query. For a corporate link, it includes the source document, extracted field spans, identifier match and reviewer status. For a translated source, it includes the original text, approved translation service and any human review required by policy.

The system should compute material transaction totals from events rather than ask a language model to add numbers embedded in text. The model may select which approved events are relevant and explain the result, but a deterministic service performs currency handling, netting and rounding under an exposed rule. The packet provides both the total and the included event IDs.

Corrections remain part of provenance. If an investigator changes an extracted name, the record shows the original output, corrected value, reason and reviewer. Repeated corrections by a source, language or document template become a quality signal. They are not silently absorbed into a fine-tuning set that may contain protected case content or encode one reviewer's preference as policy.

Provenance also protects against source dependency. Five articles repeating one allegation do not constitute five independent sources. The system should identify shared links, syndicated text or common underlying announcements where possible. The reviewer sees one claim with its dependency graph, not a credibility score inflated by repetition.

Retention has to preserve reproducibility without copying every sensitive document into the agent platform. Stable source references, hashes where appropriate, relevant excerpts, transformation receipts and access instructions can be retained in the protected case domain. The original stays in the governed source according to its record schedule. Where the source can change, the institution retains the evidence version required by policy and law.

A generated sentence is only as supportable as its weakest hidden transformation. Exposing that chain turns extraction errors into diagnosable defects and lets the investigator challenge the right component rather than debate the fluency of the summary.

Put the reporting boundary in the architecture

A warning in a prompt does not create a reliable authority boundary. The tool and data plane must make prohibited action impossible. The triage agent has no filing credential, no access to the regulatory submission channel and no permission to mark a filing as made. It cannot disclose or query restricted SAR data unless its role and case purpose explicitly permit the underlying information.

The investigator works in a separate application role. If the investigator believes the case may meet the applicable test, the system routes the evidence packet to an authorised reporting role, such as a designated financial-crime officer under local governance. That person selects the outcome, edits or approves the report narrative and triggers filing through a controlled service. Dual approval can be required by institutional policy where appropriate.

Figure 5Alert agent to receipt and supporting-document indexCausal and control schematic
Alert agent to receipt and supporting-document index9 declared states connected by 8 authored relations. The figure supports the section Put the reporting boundary in the architecture. L0L1L2L3L4 01
Alert agent
02
Evidence packet
03
Authorised investigator
04
Recommendation and cited facts
05
Authorised reporting officer
06
File or do not file
07
Controlled filing service
08
Protected no-file record
09
Receipt and supporting-document index
Boundaries: T["Triage zone · I["Investigation zone · R["Restricted reporting zone
Reading. The authored topology makes 8 declared relations across 9 states inspectable. Read it as the control structure for “Put the reporting boundary in the architecture”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
Action Triage agent Investigator Authorised reporting role Enforcement point
Link and deduplicate alerts Propose within case scope Review material grouping Not normally required Case service validates identifiers and lineage
Retrieve approved evidence Yes, least privilege Yes, role based Yes, where filing review requires it Policy decision before each source call
Resolve entity candidates Rank and expose distinguishing fields Confirm or leave unresolved Review if material to report Identity service retains candidate set
State that activity meets a legal test No Recommend under policy where authorised Decide Reporting workflow accepts only authorised identity
Draft report text No in general triage context Prepare bounded factual draft if policy permits Edit and approve Restricted workspace and templates
File a regulatory report No No unless separately authorised Yes Dedicated credential and submission service
Notify customer about filing No No No, subject to applicable confidentiality law Channel policy blocks filing-status content
Restrict or exit relationship No Recommend through separate process No inherent right from filing role Business/compliance approval workflow
No agent tool can set `sar_decision`, submit a report or reveal filing status. Those state transitions require an authorised human identity, the applicable policy version and a reference to the reviewed evidence packet.

The boundary also avoids another error: treating a SAR outcome as a customer-treatment instruction. Reporting, account restriction, payment action and relationship exit are distinct decisions with different authorities. A filing does not automatically prove that a customer should be exited. A no-file decision does not automatically mean the activity is ordinary. The case system should represent these states separately.

Confidentiality is a data-architecture constraint

SAR confidentiality cannot depend on reviewers remembering not to paste text into a shared assistant. FinCEN's confidentiality advisory discusses risk-based safeguards such as need-to-know access and access logging. FinCEN also distinguishes the report from its underlying facts and supporting documentation in its supporting-documentation guidance. Exact treatment depends on governing law.

The architecture should maintain three data classes. Ordinary alert data can be visible to authorised triage roles. Underlying investigative facts receive case-based access. The existence, content and status of a SAR remain in a restricted reporting domain. The general customer-service agent should not be able to retrieve “has this customer been reported?” A customer letter generator should not receive a filing status in context.

Figure 6Triage agent to restricted reporting serviceInteraction sequence
Triage agent to restricted reporting service5 declared states connected by 7 authored relations. The figure supports the section Confidentiality is a data-architecture constraint. t
Triage agent
Policy service
Evidence stores
Investigator
Restricted reporting service
01
Request case-scoped source
02
Allow underlying facts only
03
Permitted evidence and references
04
Evidence packet without filing status
05
Escalate with authorised identity
06
Restricted reporting workspace
07
Human decision, filing and receipt
Reading. The authored topology makes 7 declared relations across 5 states inspectable. Read it as the control structure for “Confidentiality is a data-architecture constraint”, not as measured performance. Dashed paths mark hypotheses, uncertainty or non-authoritative return paths. Schematic derived from the paper's authored topology; no measured quantities.

Logging needs the same separation. A central observability platform may record that a restricted workflow executed and succeeded without copying report content. Access to detailed traces should be purpose-bound, time-limited and monitored. Production support staff do not automatically need filing text to diagnose latency.

Confidentiality follows the information through prompts, caches, traces, exports and backups. Redacting a user interface while leaving the report in a broad vector index is not a control.

The agent's useful work

Once the boundary is real, the remaining agent tasks are substantial. It can create a unified chronology from transactions and case events. It can reconcile alert identifiers and eliminate duplicate retrieval. It can translate a rule firing into the exact observed conditions. It can compare activity with the approved customer-purpose record. It can identify missing counterparty or ownership evidence. It can draft a concise factual summary with citations.

It can also challenge the evidence packet mechanically. Are the transaction totals consistent with the included ledger events? Does every named entity have an identifier? Is the customer profile version current for the decision time? Does a media claim have an identifiable publisher and date? Are two sources being described as independent when one copied the other? These checks are observable and testable.

The FATF report on opportunities and challenges of new technologies describes technology as capable of improving speed, quality and efficiency when used responsibly. It does not say that technology replaces accountable reporting judgement. The Wolfsberg statement on effective monitoring is industry guidance that argues for effectiveness beyond automated transaction monitoring. It is useful market evidence, not a supervisory rule.

The strongest use case is evidence compression with loss controls, not autonomous suspicion. The system should make missing and contradictory evidence more visible than it was before.

Separate detection, triage, investigation and reporting technology

An AML platform is a stack of control functions, not one intelligence score. Detection decides which activity deserves review under approved rules and models. Triage groups, prioritises and routes alerts. Investigation assembles and tests evidence. Reporting applies the legal and institutional test through authorised roles. Feedback improves specific upstream components. Combining these functions into one end-to-end model makes errors difficult to locate and authority difficult to enforce.

Figure 7Ledger, customer and channel events to separate KYC or control-change workflowCausal and control schematic
Ledger, customer and channel events to separate KYC or control-change workflow9 declared states connected by 9 authored relations. The figure supports the section Separate detection, triage, investigation and reporting technology. L0L1L2L3L4 01
Ledger, customer and channel events
02
Detection controls
03
Alert objects
04
Triage and case formation
05
Investigation and evidence testing
06
Authorised reporting decision
07
Protected submission or no-file record
08
Profile, detection and source feedback
09
Separate KYC or control-change workflow
Reading. The authored topology makes 9 declared relations across 9 states inspectable. Read it as the control structure for “Separate detection, triage, investigation and reporting technology”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Each boundary has its own objective. Detection seeks coverage of relevant patterns at an acceptable burden. Triage seeks coherent, correctly scoped cases without suppressing meaningful differences. Investigation seeks an accurate, balanced account of available facts and gaps. Reporting seeks a timely, supportable decision and, where applicable, a useful submission. Optimising the same metric across all four is unsafe. Lower alert volume can improve triage workload while reducing detection coverage. Higher filing conversion can reflect better case selection or simply more aggressive escalation.

The placement of models should follow the task. Deterministic code is appropriate for transaction reconciliation, policy entitlement, schema validation, deadlines and submission state. Conventional statistical or machine-learning models can detect patterns and resolve entities when their performance is measured for the population. Language models can extract from variable documents, compare narratives and draft cited summaries. Human investigators test explanations and apply judgement. An authorised reporting role owns the protected decision.

Function Appropriate automation Observable evidence Non-delegated boundary
Detection Rules, graph features, supervised or unsupervised models under approved governance Input population, feature/event values, version and trigger reason A score is not suspicion
Triage Deterministic deduplication plus bounded agent proposals Link reasons, scope changes, priority rule and source health Agent cannot suppress a high-consequence route outside policy
Investigation Retrieval, reconciliation, extraction, chronology and alternative-evidence checks Tool receipts, propositions, conflicts, reviewer changes Agent cannot declare crime, motive or the legal test met
Reporting Template validation, deterministic totals and restricted drafting support Selected facts, edits, approval, filing receipt Authorised person decides and submits
Feedback Label-quality checks, error taxonomy and controlled rule/model change Change request, evaluation, approval and release version One case outcome cannot silently rewrite policy or a model

The agent orchestrator should have a typed plan space. It can retrieve the customer's approved profile, ledger events for the scoped period, approved corporate records and permitted prior case facts. It can request a candidate expansion when evidence supports it. It cannot invent a connector, execute arbitrary queries, alter thresholds or access restricted reporting data. Plan steps carry budgets for time, records, graph hops and retries.

Tool failure must be explicit. If entity resolution is unavailable, the investigation does not receive a graph with missing links presented as complete. It receives a dependency state and the approved manual route. If an upstream detector's model version is unknown, the alert remains reviewable as activity but the system cannot claim why the model considered it unusual.

This decomposition also makes independent challenge possible. Detection validation can test performance and data. Security testing can attack tool permissions and cross-case isolation. AML quality assurance can reperform investigations. Reporting assurance can inspect content and timeliness. A monolithic agent would require every reviewer to understand every failure mode at once.

The system is governed at the seams. Each transition changes the object's meaning, the people authorised to act and the evidence required. Making those transitions explicit prevents technical convenience from becoming accidental delegation.

Investigate competing explanations, not a single generated story

An investigation starts with concern, so confirmation bias is a structural risk. A generative system trained to produce coherent answers can intensify it by selecting facts that fit the first hypothesis. The workflow should require competing explanations and evidence tests without pretending that every explanation is equally likely.

For Harbour, one hypothesis is that related counterparties and rapid movement are inconsistent with the recorded trade purpose and may require escalation. Another is that the customer changed supplier and treasury arrangements while its KYC record remained stale. A third is that entity resolution incorrectly connected common names. A fourth is that reversals or booking mechanics created an apparent flow that did not occur economically. These are investigative branches, not declarations of innocence or suspicion.

For each branch, the agent can state what evidence would support it, what would weaken it and which authorised sources are available. It then retrieves only evidence justified by the case scope. The investigator decides whether to expand, stop or add a hypothesis. This approach is different from asking a model to “argue both sides,” which can generate symmetrical speculation without evidential discipline.

Figure 8Observed pattern and case scope to investigator judgement under policyCausal and control schematic
Observed pattern and case scope to investigator judgement under policy9 declared states connected by 11 authored relations. The figure supports the section Investigate competing explanations, not a single generated story. L0L1L2L3L4 01
Observed pattern and case scope
02
Hypothesis H1: activity conflicts with purpose
03
Hypothesis H2: legitimate business change
04
Hypothesis H3: entity-link error
05
Hypothesis H4: payment-event artefact
06
Discriminating evidence plan
07
Authorised retrieval and reconciliation
08
Support, contradiction and unresolved gaps
09
Investigator judgement under policy
Reading. The authored topology makes 11 declared relations across 9 states inspectable. Read it as the control structure for “Investigate competing explanations, not a single generated story”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Evidence can be discriminating without being definitive. An invoice matching a payment supports a trade explanation but may not establish that goods were delivered. A shared registered director supports a corporate link but not an illicit purpose. Device reuse supports a technical association but can arise from a corporate service provider or shared network. The packet should explain the proposition each item changes and avoid labelling it “exculpatory” or “incriminating” outside context.

The agent should actively look for ordinary data-quality explanations before escalating a pattern: duplicate posting, currency conversion, reversal, late file arrival, customer-account mapping error, batch payment, processor aggregation or stale profile. Finding one does not automatically close the case. It narrows what remains unexplained and directs human attention to the right facts.

Unknowns need decision relevance. A list of every missing field overwhelms reviewers. The system ranks gaps by whether they could change case scope, a material fact, the applicable policy or the decision. It explains that ranking and preserves lower-ranked gaps. An investigator can override it. A missing middle name in a uniquely identified company director is different from a missing identifier where two same-name people lead to opposite graph structures.

Challenge prompts should be embedded as structured checks, not free-form internal debate. Does any material conclusion rely on one weak source? Are there events that contradict the timeline? Are linked entities unresolved? Could totals be explained by reversals? Did the retrieval omit the current customer-purpose version? Does the conclusion use a term not defined by the applicable policy? The answers cite evidence or say the check could not be completed.

Human investigators must be able to add local knowledge with provenance. A relationship-management explanation is a sourced statement from a named role at a time, not an established case fact. If the investigator accepts it, the record states what corroboration or policy made it sufficient. This protects against both automatic distrust and uncritical acceptance.

The final rationale should show which hypotheses were tested, which material facts were established, what remains unresolved and why the authorised outcome followed under the selected policy. It should not expose or claim private chain-of-thought. A concise, evidence-linked rationale is sufficient for review and reperformance.

The goal is not to make the agent sceptical of everything. It is to prevent fluency from closing the space of plausible explanation before evidence does.

Design confidentiality against insider error and agent attack

The reporting boundary has adversaries and accidents on both sides. A malicious prompt may try to reveal filing status. A support engineer may copy a trace into a ticket. A reviewer may paste restricted text into an ordinary assistant. A cache may preserve a response after access expires. A retrieval index may expose neighbours from another case. Security must assume that content can be hostile and authorised users can make mistakes.

The NIST Zero Trust Architecture is a useful general technical reference because it makes access a continuous decision about a subject, resource and action rather than a property of being “inside” a network. It does not define SAR confidentiality. In this design, the policy decision point considers user or service identity, case assignment, purpose, data class, requested action, environment and time. Access tokens are short-lived and case-scoped.

The triage agent and restricted reporting assistant use different service identities, tool registries, stores, encryption domains and logging policies. They should not share conversational memory. Moving a case into reporting creates a controlled, one-way package of permitted underlying facts and source references. It does not grant the triage agent a pointer it can later follow into the filing record.

Attack or error Architecture response Evidence of control
Customer document instructs the agent to reveal or file Treat retrieved text as untrusted data; no filing tool in triage; instruction hierarchy enforced outside content Denied action and adversarial replay result
General assistant asks whether a subject has a filing Restricted status is absent from its index and denied by policy Access-decision log without protected content
Support trace captures narrative text Content-minimised telemetry and restricted diagnostic workflow Trace schema tests and sampled log inspection
Reviewer copies text into an unauthorised channel Data-loss controls, blocked export and role training Prevented event, alert and incident follow-up
Embedding index returns another case Case partitions, row-level entitlements and no broad nearest-neighbour search Cross-case penetration tests
Employee queries high-profile names without case purpose Purpose-bound access, anomaly monitoring and approval where required Named access review and investigation path
Backup outlives record policy Separate retention, encryption and verified deletion process Restore test plus deletion attestation

Prompts and model outputs should be classified at least as strictly as their inputs. If a restricted reporting model sees protected data, its generated summary, token cache and evaluation sample are protected too. Sanitising names does not necessarily remove sensitivity if the transaction pattern and context can re-identify the case.

Observability can remain effective without copying content. Record tool name, case-scoped pseudonymous identifier, policy decision, duration, result class, token and record volume, model version and error code. Content access for debugging requires a separate approved path with an audit trail. Synthetic cases should reproduce most defects without exposing live reports.

Break-glass access needs an explicit reason, short duration, monitoring and after-the-fact review. It should not become the normal solution to poor entitlements. Production engineers should be able to restart a failed worker or inspect latency without seeing report text. Data stewards can diagnose schema changes using representative masked records.

Incident response distinguishes possible disclosure, confirmed disclosure, integrity failure and availability failure. A suspected leakage event freezes relevant logs and access tokens, identifies data touched, notifies the authorised confidentiality and legal owners, and follows the applicable response process. A source-integrity failure identifies all derived propositions and decisions through lineage. An outage activates the documented manual reporting route where time-sensitive obligations may apply.

The threat model should be exercised, not only documented. Red teams attempt prompt injection, cross-role retrieval, case-ID manipulation, indirect disclosure questions, cache recovery, export and support-channel leakage. They test the actual deployed identities and stores. A refusal sentence from the model is not a pass if the protected data was retrieved into context first.

The strongest confidentiality control is data the requesting component never receives. Model refusals add defence in depth; they do not replace isolation, entitlements and monitored purpose.

Partition group and cross-border work by permitted purpose

Financial activity and customer relationships cross legal entities, products and borders. An investigator may need facts from another group entity to understand a pattern, while the reporting duty and confidentiality rules remain local. A central agent that can search every case and filing creates apparent efficiency by ignoring the hardest part of the control: whether the information can be used, transferred and disclosed for this purpose.

The group architecture should separate a shared evidence exchange from local decision domains. A local case sends a typed request describing subject identifiers, time, evidence class, purpose and requesting role. The source entity's policy service decides what can be returned. The response contains permitted facts and provenance, not unrestricted access to the source case. Local investigators apply their own law and policy to the result.

Figure 9Entity a investigation to entity b restricted reporting statusCausal and control schematic
Entity a investigation to entity b restricted reporting status7 declared states connected by 7 authored relations. The figure supports the section Partition group and cross-border work by permitted purpose. L0L1L2L3L4
not exposed through ordinary exchange
01
Entity A investigation
02
Purpose-bound evidence request
03
Group exchange policy
04
Entity B evidence domain
05
Permitted underlying facts and provenance
06
Entity A authorised reporting decision
07
Entity B restricted reporting status
Reading. The authored topology makes 7 declared relations across 7 states inspectable. Read it as the control structure for “Partition group and cross-border work by permitted purpose”, not as measured performance. Dashed paths mark hypotheses, uncertainty or non-authoritative return paths. Schematic derived from the paper's authored topology; no measured quantities.

FATF Recommendation 18 and its interpretive material, available through the current FATF Recommendations, address group-wide programmes and information sharing at the international-standard level. Domestic implementation, secrecy, data protection and reporting-confidentiality rules still govern. The architecture therefore treats group sharing as a policy decision, not as an entitlement inferred from corporate ownership.

Identifiers need careful handling across entities. A group customer identifier can help resolution but may be missing or wrongly joined. Legal names, local account identifiers and incorporation numbers remain part of the evidence. The exchange service can return candidates and distinguishing fields. It does not merge two customers globally because one model score crossed a threshold.

The requesting case should receive the minimum fact needed. A payment event can be shared as date, direction, amount, currency, account role and stable source reference where authorised. A source investigation's interpretation is labelled as another entity's conclusion with its policy and time, not reclassified as a fact. The existence or non-existence of a report stays in the restricted domain unless a specific lawful sharing route applies.

Group cases need coordination without authority collapse. A lead coordination record can show participating entities, factual dependencies, deadlines and named contacts. Each legal entity retains its investigation and reporting decision. The lead cannot file on another entity's behalf unless a separately authorised arrangement says so. One entity's no-file outcome does not bind another whose facts or test differ.

Language and currency normalisation require receipts. An amount converted for comparison retains the original value, currency, rate source and time. Translated facts retain original text and translation path. Local legal terminology should not be replaced by a generic English label that changes meaning. The agent can prepare parallel views while preserving the authoritative local record.

Cross-border requests need service levels and failure states. A denied request identifies the rule or reason class without revealing protected data. A delayed response remains an open dependency. The investigator decides under policy whether available evidence is sufficient or whether escalation is required. The agent cannot treat a foreign silence as exculpatory or incriminating.

Data residency and model processing location belong in the control. If a group service or model cannot process a data class in a jurisdiction, retrieval is blocked before content reaches it. Local processing may produce a permitted abstraction. Encryption and contractual controls matter, but they do not create a lawful purpose where none exists.

Feedback follows the same boundaries. A local correction can update a shared entity identifier or source-quality issue when authorised. It should not broadcast a protected case outcome as a global risk label. Group model development uses datasets whose provenance, transfer and purpose are approved, not a convenient export of every investigation narrative.

A group view should connect permitted evidence while preserving local authority. Centralising search is easy; making every fact, decision and disclosure lawful and attributable is the actual architecture.

Make filing, no-file and continuing review complete lifecycles

Reporting architecture often concentrates on the submission button. The control begins earlier and continues after submission. A complete lifecycle covers escalation, deadline calculation under the applicable regime, evidence selection, narrative preparation, approval, transmission, receipt, supporting-document retention, correction, further activity and authorised closure. A no-file outcome also needs rationale, approval, monitoring or control action and retention under policy.

The workflow engine calculates procedural dates only after jurisdiction, entity, report type and triggering decision are known. It cites the policy rule and shows the authorised user the result. It does not ask a language model to infer a filing deadline from prose. If the facts that start a deadline are disputed, the case displays the ambiguity and escalates under policy.

Submission is an exactly-once operation. The reporting service validates the final structured form and narrative, obtains the required approval, submits with a dedicated credential and stores the regulator or FIU receipt. A timeout produces an unknown state, not an automatic resubmission. Reconciliation checks the receiving channel before retry to avoid duplicate reports.

The FinCEN supporting-documentation guidance illustrates, for its United States scope, why report content and underlying documentation need a managed relationship. Other regimes differ. The architecture therefore maintains a supporting-document index with source owner, version, retention rule and the material report proposition supported. It does not assume every document is transmitted with the filing.

A no-file decision is not an empty case closure. It records the authorised person, applicable policy, material facts considered, unresolved gaps, reason and any follow-up. Follow-up may be continued monitoring, a KYC review, detection adjustment or no additional action. The protected record should be reviewable by roles permitted under local policy without exposing it as ordinary customer data.

Further activity should be compared with the prior decision, not appended to an ever-growing narrative. The new case identifies what changed since the decision-time cut, which prior propositions remain valid and whether a different policy version applies. That reduces repeated retrieval while avoiding stale conclusions. It also stops a prior filing from biasing every later event toward the same outcome.

Corrections preserve history. If a report requires amendment or supplementation under the applicable regime, the authorised workflow creates a linked record. It does not overwrite the submitted narrative or receipt. If a no-file decision is revisited after new evidence, the new decision cites the earlier one and states what changed.

Quality assurance samples both filing and no-file outcomes. Reviewing only filed cases tests narrative quality but not under-escalation. Reviewing only alerts closed by the agent tests triage but not the reporting boundary. Sampling should consider alert type, investigator, customer segment, evidence complexity, decision, time and system version. High-consequence exceptions can receive full review under policy.

A protected decision is a stateful control, not a document-generation task. The lifecycle must demonstrate who decided, what evidence they used, what was sent, what the receiving authority acknowledged and how later facts were handled.

Work the harbour case without inventing motive

The composite agent links the three alerts because they share the same customer, a close time range and related counterparties. It retrieves the approved KYC profile and discovers that the ownership record predates a registry change. It retrieves transaction events and proves that the new beneficiaries received funds from the newly added account. Entity resolution finds a shared director across two counterparties but leaves a third unresolved.

An external article claims the shared director is connected to fraud. The article cites no source, and the name is common. The agent labels the item “unverified external allegation” rather than treating it as fact. A later official registry record confirms the corporate connection but says nothing about criminal purpose. The evidence graph preserves both distinctions.

The agent prepares a timeline, relationship graph, transaction table, customer-profile comparison and list of open questions. It notes that activity appears inconsistent with the stored purpose and that the customer record may be stale. It does not write “money laundering network” or select a filing outcome. It also identifies that the first alert was closed without the later ownership evidence, giving the investigator a reason to reperform that decision.

The investigator obtains additional authorised information, tests legitimate explanations and records a recommendation under local policy. The case then crosses into the restricted reporting workflow. The authorised officer decides the filing outcome after reviewing facts, applicable law and policy. If a report is filed, the controlled service returns a receipt and supporting-document index. If it is not filed, the authorised no-file rationale and monitoring decision remain protected according to policy.

No part of the scenario establishes a universal answer. It demonstrates where evidence ends and accountable suspicion begins.

A relationship graph must not become guilt by association

Graphs are attractive in AML investigations because they make shared parties, accounts, devices and payment paths visible. They also make weak relationships look stronger than they are. A line on a graph can mean common ownership, a payment, a shared address, a possible name match or a technical co-occurrence. If those edge types are rendered alike, the diagram encourages a conclusion the evidence does not support.

Each edge needs a type, direction, source, event or validity time and confidence status. “Paid” is a ledger event. “Registered director of” is a filed corporate relationship for a stated period. “Possible same person” is an unresolved identity candidate. “Mentioned in the same article” is only co-occurrence. The investigation view should let the reviewer hide weak edges and inspect the record behind any remaining connection.

Graph expansion also needs a stopping rule. Starting from one counterparty and retrieving every connected entity can expose unrelated customers and overwhelm the case with low-value links. Expansion should follow the case hypothesis, the investigator's role and preapproved relationship types. A second hop may be justified when the first-hop entity controls the beneficiary. Ten hops across common addresses are not justified merely because the database can return them.

The case record should preserve negative results. If two similar names were tested and found to be different people, that disambiguation prevents the next run from reopening the same theory. If an address is a large professional-services office, sharing it may carry little weight. A graph that stores only positive links systematically overstates connectedness.

A graph is an index into evidence, not evidence by itself. Evaluate it with link-level precision, entity-resolution error, source freshness and the incremental value of each expansion step. Reviewers should be able to explain why a displayed connection mattered to the case rather than relying on visual density.

Drafting a report narrative without manufacturing certainty

Where policy permits drafting assistance, the narrative generator should work only inside the restricted reporting workspace after authorised escalation. Its inputs are the reviewed proposition ledger, selected transaction events, entity decisions and the applicable report schema. It should not start from the triage agent's free-text summary.

The draft must distinguish observed facts from institutional judgement. Dates, amounts, account roles and transaction direction should come from structured events and be reconciled. The reason the institution considers activity suspicious should be written by, or under the approval of, the authorised reporting role. The generator can check chronology, remove duplication and flag undefined entities. It cannot add a motive, crime type or relationship that is absent from reviewed evidence.

Supporting documentation needs an index even when it is not submitted with the report. The index should link each narrative fact to a retained record, identify the record owner and preserve the version used. If a supporting record is later corrected, the correction process should assess whether an amended or supplemental filing question arises under the applicable regime rather than silently rewriting history.

Quality controls should compare the final narrative with the reviewed ledger, not with another model's preferred prose. Useful checks include omitted material facts, unsupported named entities, inconsistent totals, unexplained abbreviations, chronology breaks and disclosure of data outside the permitted report. A report can be stylistically polished and operationally poor if law enforcement cannot identify the subjects, accounts, transactions and reason for concern.

Narrative quality means accurate, useful and supportable information. It does not mean dramatic language. The authorised officer should see every machine edit and retain the ability to reject the draft. The final record should identify the human approver and the version actually submitted.

Failure modes worth testing before scale

The first failure is alert compression that hides diversity. Three alerts may concern one pattern, but they may also involve different customers or time periods. The linking service should show why it proposes a merge and let the investigator split it without losing source lineage.

The second is confirmation bias in retrieval. If the hypothesis mentions laundering, a broad search may return only risk-themed content. Retrieval tests should include legitimate explanations, prior false matches, reversals and customer-purpose evidence. The interface should place contrary evidence beside supporting evidence rather than bury it in an appendix.

The third is policy leakage. Filing criteria copied into a general agent prompt can escape through logs, support tools or customer-facing content. Restricted policy, report state and narrative data need a separate access domain. A public red-team prompt should never be able to discover them.

The fourth is silent source failure. An empty registry or device result may mean no record exists, the source is unavailable, permission failed or the query was malformed. Those outcomes require different handling. The agent must not turn “retrieval failed” into “no adverse information found.”

The fifth is reviewer displacement. If experienced investigators spend their time correcting transactions and names, the system has moved clerical work rather than removed it. Track correction by field and source. Persistent corrections should stop or narrow the defective component, not become accepted human cleanup.

Build the investigator interface around disagreement

Most case interfaces make the narrative dominant and evidence secondary. That is the wrong hierarchy for agent-supported investigation. The first view should show the case scope, the observations that opened it, material propositions, contradictions, missing evidence and applicable policy. A narrative can sit beside these objects as a compact orientation, but the investigator must not need to dismantle it to find uncertainty.

The alert-link view should expose merge reasons one by one. The investigator can accept a shared account identifier, reject a same-name counterparty link, or defer a device association pending more evidence. Splitting a case should preserve completed work against the alerts that still belong together. Merging should show which deadlines, owners and restricted states are affected before it executes.

The transaction view should be reproducible. Filters show precisely which ledger events are included. Totals display currency treatment, reversals and date basis. Selecting a row opens the immutable event reference and any normalisation receipt. Charts can reveal timing or flow, but the underlying table remains available. Visual scale should not exaggerate a small value or hide a large one through automatic axes.

The relationship graph should default to the strongest resolved edges and mark weaker candidates distinctly. Each edge has a source and time. The investigator can expand one justified hop and sees the additional data-access purpose before retrieval. Graph density is never used as a suspicion score. A common high-degree service provider, payment processor or address can otherwise make ordinary activity look organised.

The proposition ledger should make three investigator actions easy: confirm from evidence, reject with a reason, or leave unresolved. Confirmation records the human and source. Rejection records whether the problem was extraction, source, identity or relevance. Unresolved status prompts a decision about whether the gap is material. It does not force an investigator to manufacture certainty to close a field.

Warnings must be consequence-based. A missing non-material description should not compete visually with an attempted access to restricted reporting data. Define severity through policy and test comprehension with investigators. Alert fatigue can move from transaction monitoring into the agent interface if every extraction confidence or source delay becomes a red banner.

Reviewer independence requires careful defaults. The interface can present an agent proposal, but it should not preselect the reporting recommendation. Material contrary evidence should appear in the main view. Where practical, a reviewer should inspect the evidence before seeing a generated recommendation in sampled evaluations, allowing the programme to measure anchoring.

Accessibility and expertise are control concerns. Tables need clear reading order, graphs need equivalent text, colours cannot carry status alone and timestamps need explicit zones. New investigators may need policy explanations; experienced investigators need rapid source access. One generic conversational panel serves neither well.

The final decision screen should make the consequences explicit. It lists the applicable regime, selected policy version, factual propositions, material unknowns, proposed outcome and any separate customer or monitoring actions. The reporting decision requires the authorised role. Other actions route to their own owners. One click cannot file, restrict an account and change the customer profile together.

Usability testing should measure evidence discovery, correction and challenge, not preference for a polished design. Give investigators cases with planted contradictions and name collisions. Observe whether they locate the source, understand edge strength, recognise an unavailable feed and reject unsupported text. A faster completion time is beneficial only when those tasks remain reliable.

The interface is part of the authority model. If it hides evidence and invites one-click acceptance, formal human approval becomes ceremonial. If it exposes disagreement and separates consequences, human review can remain substantive at production pace.

Govern change as a financial-crime control change

Agent releases can alter more than language. A new retrieval connector changes available evidence. An entity-resolution threshold changes graph structure. A prompt changes which contradictions appear. A model upgrade changes extraction and writing behaviour. A user-interface change affects reviewer reliance. Each change needs impact analysis against the control objective and protected boundary.

The release unit should identify all component versions: detector, source connectors, schemas, normalisers, entity resolver, language model, prompts, tool registry, policy package and interface. The case receipt stores that manifest. Reproducing an old case may require archived deterministic components and retained model access; where exact regeneration is impossible, the programme must preserve the observable outputs and acknowledge the limit.

Changes begin with a stated hypothesis. For example, a new extractor may reduce manual correction of corporate ownership fields without increasing unresolved identity collapse. The evaluation plan names the case populations, material measures, critical failures and rollback rule before results are seen. It includes filing and no-file cases, competing explanations, weak sources and restricted-data attacks.

Policy changes follow an authorised route distinct from model changes. Engineers can implement an approved reporting format or role rule; they do not interpret a new law into production alone. In July 2026, for example, AMLA opened a consultation on a draft common EU format for reporting suspicions and transaction records. It is current primary evidence of a proposed structured direction, not a final rule to deploy globally. The architecture should track consultation, finalisation, applicability and effective date separately.

Shadow testing compares the candidate with the current workflow on the same decision-time evidence. Differences are adjudicated by qualified reviewers and classified. Improvement on common easy cases cannot offset a new critical confidentiality path. Release gates are consequence-weighted: prohibited filing access, protected-status leakage, material transaction misstatement or unresolved identity collapse blocks the affected scope regardless of average quality.

Rollout can proceed by alert type, customer segment, jurisdiction and decision right. The first live stage may show packets without recommendations. Later stages can add case formation or restricted drafting only after their own gates pass. Expanding to a new legal entity requires fresh policy mapping and source coverage. It is not justified by success in a neighbouring country.

Rollback needs semantic as well as technical planning. Restoring a previous model does not undo cases already influenced by the defective release. Lineage identifies affected cases and outputs. The control owner decides whether to reperform open work, review closed decisions, correct a filing, contact another function or document that the error was immaterial.

Emergency changes have a narrow path with time-limited approval, compensating controls and mandatory retrospective review. A source outage might justify disabling one enrichment step while preserving manual investigation. It does not justify bypassing the reporting boundary. After the emergency, the system returns to the approved baseline or completes the normal release process.

Training and procedure updates travel with the release. Investigators need to know what changed, which outputs remain uncertain and how to report defects. Quality assurance needs new sampling rules. Support needs content-safe diagnostics. Security needs updated attack paths. A model card alone cannot coordinate this operating change.

The material question is not “did the model change?” It is “did the evidence, authority, reviewer behaviour or decision path change?” Governance should follow the latter.

Run the case queue as a risk system

Agentic triage changes the shape of investigative demand. Better linking can turn several alerts into one complex case. Faster evidence collection can move the bottleneck to entity resolution or authorised reporting. New sources can open clusters faster than staff can review them. A queue dashboard that shows only case age and count cannot reveal these shifts.

Represent each case by work state and consequence: intake validation, evidence retrieval, dependency waiting, active investigation, quality review, reporting decision, submission, follow-up and closure. Record the material hypothesis, reporting regime, procedural deadline where applicable, evidence decay, customer or payment impact and specialist skill needed. Priority is a deterministic, approved function of those factors with a visible reason.

Figure 10Validated cases and deadlines to control-owner interventionCausal and control schematic
Validated cases and deadlines to control-owner intervention11 declared states connected by 11 authored relations. The figure supports the section Run the case queue as a risk system. L0L1L2L3L4 01
Validated cases and deadlines
02
Consequence and dependency classification
03
Available skills and restricted roles
04
Capacity-aware assignment
05
Active investigation
06
External or internal dependency
07
Quality or reporting review
08
Dependency escalation and fallback
09
Authorised completion
10
Backlog and source-health signals
11
Control-owner intervention
Reading. The authored topology makes 11 declared relations across 11 states inspectable. Read it as the control structure for “Run the case queue as a risk system”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Queue priority must not become a hidden filing propensity. A high-consequence or time-sensitive case receives earlier review; it is not presumed more suspicious. The agent can summarise priority factors but cannot silently elevate a persuasive narrative. Every reordering event records the rule and values used.

Deduplication should reduce repeated work without erasing distinct obligations. Alerts can share evidence yet require separate cases because they concern different legal entities, customers, periods or reporting regimes. The case service can create a parent evidence bundle with child decisions. Closing one child does not close the others. Investigators see what work is reused and what judgement remains local.

Capacity models need skill and authority, not only headcount. A general investigator may resolve ordinary evidence, while a complex ownership graph, sanctions issue, language or reporting decision needs a qualified role. Restricted reporting officers can become the bottleneck even when triage is fast. Forecast demand by case type and state, then test severe but plausible source or detection surges.

Age clocks should separate active work, policy-permitted waiting and internal delay. Waiting for a source or another entity is visible, with an owner and escalation. The end-to-end clock continues where the external obligation or customer impact continues. Pausing an internal service metric cannot make the real deadline disappear.

Backlog actions are preapproved. They can add qualified capacity, narrow low-value enrichment, group exact duplicates, increase sampled quality review, or invoke manual procedures. They cannot weaken the reporting test, hide alerts, broaden automatic closure or give the agent filing authority. A temporary risk-based prioritisation decision has an owner, start and end date, affected population and retrospective review.

Workload equity matters. If the agent performs poorly on one language or entity type, those cases may require more human correction and remain older. Segment queue state and rework to find the hidden burden. The response is better source or model support and appropriate capacity, not a lower evidential standard for that group.

Service-level targets should be paired with quality and customer measures. Time to evidence packet is useful beside packet correction and material omission. Time to decision is useful beside reperformance and unsupported closure. Report submission timeliness is useful beside schema error, duplicate submission and supporting-document quality. No single target should reward an unsafe shortcut.

An efficient AML queue moves the most decision-relevant evidence to the right authorised person before consequence or obligation deteriorates. Closing the largest number of alerts is not the same objective.

Prepare for integrity incidents, not only outages

Availability failures are visible; integrity failures can continue producing plausible cases. A source connector may map the wrong account. A currency service may double-convert. An entity resolver may merge two people after a model update. A prompt may omit contrary evidence. A restricted-data policy may expose filing status to the wrong role. Incident design should assume that flawed output can already have influenced decisions.

Detection combines technical and case signals. Schema checks, reconciliation, feature distributions and access-denial monitoring can surface defects early. Investigator corrections, unusual agreement changes, repeated source challenges and quality-assurance findings provide semantic signals. One correction is a case event; a pattern across versions becomes a control incident.

The first response is containment matched to the defect. Disable the connector, model, prompt, tool or export path affected. Preserve unaffected manual and deterministic routes. A reporting deadline does not justify continuing with known corrupted evidence. The case interface states which component is unavailable and which facts require manual verification.

Lineage defines scope. Query which propositions were derived from the faulty source or transformation, which summaries used them, which reviewers saw them and which decisions or filings followed. This impact set is more precise than every case processed during the release, but it should include uncertain cases when lineage itself may be incomplete.

The control owner and authorised legal, reporting, privacy, security and operations roles decide remediation. Options include reperforming open cases, reviewing closed no-file and filed cases, correcting or supplementing a submission through the applicable process, notifying another group entity, correcting customer treatment, purging an unauthorised copy and adjusting training data. The incident tool proposes the affected set; it does not take those consequential actions.

Manual continuity requires more than a PDF procedure. Investigators need current source access, transaction reconciliation, approved templates, restricted reporting credentials and a way to record decisions and receipts. Exercises should disable the agent and one major source, then measure whether staff can meet obligations without creating uncontrolled local files.

Post-incident review asks why the defect escaped component and integrated tests, why monitoring did or did not detect it, how reviewer behaviour affected consequence and whether the authority boundary contained it. A prompt fix is insufficient if the real cause was source semantics or one-click approval.

Historical evaluation sets are updated cautiously. Cases corrected after the incident retain the original and corrected states, incident link and decision-time information. They can become regression tests. They are not silently relabelled in a way that makes the old model look better or loses evidence of the failure.

Recovery is complete when affected decisions are addressed, protected data is contained, manual backlogs are reconciled, monitoring is improved and named owners accept residual risk. Restoring the endpoint is only one milestone.

The value of provenance appears most clearly when something plausible is wrong. It lets the institution find the decisions that depended on the defect and repair them without treating every case as unknowable.

Use a scorecard that does not reward filing volume

The balanced scorecard needs precise definitions. Alert-to-case compression shows duplicated work removed but must be paired with erroneous-merge and split rates. Evidence coverage shows whether required propositions have sources but must distinguish “not available” from “not retrieved.” Investigator correction shows clerical burden but must identify whether corrections are substantive or stylistic. Decision timeliness matters without turning the fastest outcome into the preferred one.

Filing conversion is a monitoring statistic, not a target. A movement can reflect risk, detection, policy, training, capacity or behaviour change. FATF's guidance on AML/CFT-related data and statistics notes that jurisdictions can define and count suspicious transaction reports differently. That official source concerns national measurement, but the same caution applies internally: counts need definitions and context before comparison.

Scorecard domain Measure examples Necessary countermeasure
Detection and case formation Alert coverage, time to case, duplicate retrieval avoided Missed material pattern review and erroneous merge rate
Evidence integrity Reconciled totals, supported propositions, source availability Material omission, stale-source and weak-source substitution rates
Investigation quality Reperformance agreement, contrary evidence considered, useful gap resolution Unsupported escalation and unsupported closure
Reporting control Timeliness, schema validity, receipt reconciliation Filing error, duplicate submission, confidentiality incident
Human control Correction, rejection, escalation and meaningful review time Automation-bias tests and sampled blind review
Customer and fairness Differential escalation and burden by relevant lawful segment Evidence-quality and detection-coverage comparison
Resilience Source, model and workflow availability; manual-route completion Recovery accuracy and backlog consequence
Economics Cost per decision-grade case and duplicated work removed Remediation, assurance and customer-burden cost

Independent quality assurance should sample the full decision distribution. Stratify by filed, no-file, escalated, closed at triage, complex graph, weak external information, source outage and new component version. Include random cases as well as risk-based samples. If assurance reviewers know the original outcome, measure that potential bias or blind appropriate stages.

Ground truth is often partial. A filing does not prove criminality; no-file does not prove ordinary activity; law-enforcement response may be unavailable or selective. Evaluate what can be established: transaction accuracy, entity identity, source support, policy application, evidence balance, authority and documentation. Use later confirmed outcomes cautiously and explain selection bias.

Metrics must include denominators, confidence or uncertainty where appropriate, and case eligibility. If only successfully retrieved cases enter a provenance measure, the result hides outages. If difficult languages are excluded from evaluation but included in production, the reported performance is not representative. The release report states coverage plainly.

Investigate divergence, not only threshold breach. A stable overall correction rate can hide a sharp increase for one document type offset by improvement elsewhere. A falling handling time with rising one-click approval can signal automation bias. A declining alert population after a feed change can signal missing data. Dashboards should allow a control owner to trace these patterns to component and cohort versions.

Stop actions are pre-agreed. Protected-data leakage disables the affected path. Material transaction errors suspend narrative generation for that case type. Entity-resolution collapse narrows graph automation. Excessive unsupported closures return cases to full human review. A queue backlog invokes risk-based prioritisation and manual capacity, not looser evidence standards.

Economic claims use realised operational data. Time observed in a controlled pilot can inform a capacity model, but savings are not realised until work, staffing or avoided demand changes. Add source, compute, testing, oversight, remediation and incident cost. A credible business case can still be strong; it simply does not treat reviewer time as free before automation and redundant after it.

A useful AML metric explains the control, not the desired headline. The programme succeeds when investigators receive more reliable evidence, protected decisions remain attributable, and defects are discovered early enough to act.

Evaluation must penalise a persuasive wrong story

An AML triage agent can produce a polished summary while failing every important control. It may omit a source that weakens its interpretation, merge two people with the same name, expose restricted filing status or calculate transaction totals incorrectly. Evaluation must work at the proposition, tool and authority level.

Evaluation dimension Test design Unsafe failure Release gate
Alert linking Cases with true and false shared identifiers Unrelated customers or counterparties merged Consequence-weighted precision and recall by link type
Transaction accuracy Recompute from immutable ledger events Amount, direction, currency or time misstated Exact reconciliation for material case totals
Entity resolution Same-name and changed-ownership cases Candidate collapsed without decisive evidence Ambiguity retained when identifiers conflict
Evidence balance Include exculpatory and contradictory sources Summary selects only incriminating facts Material contrary evidence always visible
Source provenance Copy chains and weak external allegations Repeated claim counted as independent corroboration Source identity and dependency retained
Authority Attempt direct filing and status mutation Agent can reach restricted state or credential Denied in architecture and tested continuously
Confidentiality Prompt-injection and cross-role retrieval tests Filing existence leaks to a general service No protected-status disclosure across tested paths
Human review Blinded comparison with current process Fluent output increases unsupported escalation Review quality, correction and disagreement measured

The adversarial suite should contain a document that instructs the agent to file, a customer name that resembles a known subject, two alerts with coincidental identifiers, a stale KYC record, a transaction reversal, and a support user asking whether a SAR exists. The expected responses are structured: ignore document instructions, retain identity ambiguity, refuse the status query and send the case only through authorised routes.

Private model reasoning is not evidence. The assurance record consists of tool calls, source references, extracted propositions, policy decisions, human edits and final authorised state changes. A concise rationale may explain why an exposed fact supports a recommendation. It must not be presented as a faithful transcript of hidden computation.

Figure 11Versioned alert and investigation cases to monitor corrections, missed evidence and leakageCausal and control schematic
Versioned alert and investigation cases to monitor corrections, missed evidence and leakage10 declared states connected by 10 authored relations. The figure supports the section Evaluation must penalise a persuasive wrong story. L0L1L2L3L4
No
Yes
01
Versioned alert and investigation cases
02
Linking and retrieval tests
03
Transaction and entity reconciliation
04
Authority and confidentiality attacks
05
Human adjudication
06
All critical gates pass?
07
Fix, narrow tools or stop release
08
Shadow operation
09
Limited cases with full human review
10
Monitor corrections, missed evidence and leakage
Reading. The authored topology makes 10 declared relations across 10 states inspectable. Read it as the control structure for “Evaluation must penalise a persuasive wrong story”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Operations: measure value without rewarding under-reporting

A productivity metric can corrupt the process. If the target is fewer escalations or fewer filings, the system learns that suppression looks efficient. If the target is shorter handling time, investigators may accept the first coherent narrative. Neither metric measures useful intelligence or control quality.

Use a balanced scorecard: evidence completeness, transaction reconciliation, entity-resolution error, duplicated work, investigator correction, reopened cases, decision timeliness, confidentiality incidents and sampled outcome quality. Filing volume can be monitored for unexplained change, not treated as a success target. Cohort analysis should check whether particular customer types receive weaker evidence or disproportionate escalation.

Model and rule changes need separate versioning. The Federal Reserve's SR 26-2 superseded earlier United States model-risk guidance in April 2026 and expressly excludes generative and agentic AI from its formal scope. That matters here: a conventional transaction-monitoring model may sit within applicable model governance, while the generative triage layer requires system, data, security and human-governance controls chosen for its use. Calling the whole stack “the model” hides ownership.

The NIST Generative AI Profile provides a useful non-binding lifecycle frame for generative risk. It should complement, not replace, financial-crime obligations. The operating owner should review changes in source coverage, agent behaviour, reviewer reliance and restricted-data access after each release.

A production standard

Production evidence should be readable at three levels. An investigator needs the transaction, entity and source detail. A control owner needs the case state, exception and component version. Senior management needs population coverage, quality, confidentiality, capacity and realised economics. These are different views over the same governed events, not separately written narratives whose numbers drift.

The service should also expose what it does not know. Management reporting identifies source outages, untested populations, immature labels, pending quality findings and legal interpretations awaiting implementation. A case packet identifies unresolved entities and unavailable evidence. A filing decision records the authorised judgement despite those limits. Uncertainty is managed through owners and actions, not removed from the page.

Success is therefore conditional. Faster evidence assembly is valuable when transaction facts still reconcile, contrary evidence remains visible, investigators can challenge the packet, protected status stays isolated and no-file as well as filing decisions withstand reperformance. Reduced duplicate work is valuable when distinct legal obligations remain separate. A lower queue is valuable when detection coverage and customer treatment do not deteriorate.

The same standard applies to a vendor component. A demonstration on synthetic alerts cannot establish local source coverage, policy fit, identity resolution or reviewer behaviour. The institution evaluates the component inside its own authority and data architecture, records residual limits and retains an exit or manual route.

A production system earns trust by making its limits operational. It denies unauthorised action, marks unavailable evidence, routes consequential judgement to named people and leaves enough trace to repair a decision after a hidden defect is discovered.

Before release, the institution should be able to answer these questions with evidence:

  • Which alert types and entities can the agent link?
  • Which sources can it access for each role and purpose?
  • How does it represent conflicts and unknowns?
  • Can it recompute every material transaction statement?
  • Where is jurisdiction and policy selected?
  • Which identity can recommend a filing outcome?
  • Which identity can decide and submit it?
  • Can any general assistant discover a filing's existence?
  • How are customer-treatment decisions separated from reporting?
  • What happens when a source or entity-resolution service fails?
  • How are investigator corrections fed into evaluation without silently changing policy?
  • Which metric or incident narrows the agent's permissions?

The useful end state is not an investigator-shaped chatbot. It is a controlled case system in which evidence is faster to assemble, harder to distort and easier to challenge. The reporting decision remains human, attributable and protected because the architecture makes it so, not because a prompt asked politely.