Home · Writing · Architecture

Context Compilers for Policy-Filtered Retrieval

A reference architecture for compiling identity, purpose, policy, time and evidence requirements into a bounded retrieval plan before an agent sees enterprise context.

TLDR

  1. A reference architecture for compiling identity, purpose, policy, time and evidence requirements into a bounded retrieval plan before an agent sees enterprise context.
  2. This article develops that architecture through a composite bank case. “Northstar Bank” and its records are fictional.
  3. The words in Maya’s question reveal almost none of this. “The customer” must resolve to a canonical party and facility.
  4. Prompt construction is a presentation concern. Compilation is a control concern. A prompt can say “only use policies that apply to India,” but that instruction is interpreted by the same probabilistic component that is expected to read the retrieved material.
  5. RFC 8693 defines OAuth token exchange patterns that can represent delegation and impersonation semantics. RFC 9396 adds structured authorization_details for fine-grained requests.
Figure 1Typed decision request to model context plus evidence manifestCausal and control schematic
Typed decision request to model context plus evidence manifest11 declared states connected by 10 authored relations. The figure supports the section Retrieval is an authorisation event. L0L1L2L3L4
No
Yes
01
Typed decision request
02
Context compiler
03
Verified principal and delegation
04
Approved policy bundle
05
Effective time and jurisdiction
06
Plan valid?
07
Deny or route to human
08
Bounded retrieval plan
09
Lexical, vector and graph retrieval
10
Post-retrieval evidence validation
11
Model context plus evidence manifest
Reading. The authored topology makes 10 declared relations across 11 states inspectable. Read it as the control structure for “Retrieval is an authorisation event”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
On this page

Retrieval is an authorisation event

A banking assistant receives a seemingly ordinary question: “Can the small-business customer keep the existing overdraft while its annual review is open?” The answer may depend on the current credit policy, the customer’s approved limit, an exception recorded by a credit officer, a product guide, a covenant, and the legal entity that booked the facility. A semantic search can find text about overdrafts. It cannot decide which customer records this employee may see, which policy applies to this entity, whether the exception is still valid, or whether a superseded guide may be used for a live decision.

The usual retrieval-augmented design begins too late. It embeds a query, fetches similar chunks, reranks them and asks a model to answer with citations. Security is represented as a metadata predicate assembled somewhere near the vector query. Policy selection is often left to prompt wording. Freshness becomes a recency boost. The result can be relevant yet unauthorised, current yet inapplicable, or well cited yet unsupported.

A production agent should not search first and justify later. It should compile a permitted evidence space before retrieval begins. The compiler is a deterministic control plane. It accepts a typed request, a verified principal, an approved purpose, policy versions and evidence requirements. It emits a bounded retrieval plan. Search services execute that plan; they do not broaden it. The model receives only the admitted evidence and a manifest that explains why each item entered the context.

This article develops that architecture through a composite bank case. “Northstar Bank” and its records are fictional. The regulatory and technical sources are public; the operating numbers are test thresholds, not reported results from a deployment.

A context compiler converts a decision request into four executable artefacts: an authorisation predicate, an applicability predicate, an evidence contract and an output budget. Its product is not prose. It is a signed, versioned plan that can be tested, replayed and denied.

The northstar question has five hidden dimensions

Maya is a commercial-credit analyst in the composite case. She works for Northstar Bank’s Indian legal entity and supports a portfolio of small manufacturers. She asks the assistant about Atlas Fasteners, a customer whose annual review is incomplete. Atlas has an overdraft, a temporary covenant waiver and a pending change in beneficial ownership. The relationship manager has added a confidential note. A group credit manual was replaced last week, while a local addendum remains effective until the end of the quarter.

The words in Maya’s question reveal almost none of this. “The customer” must resolve to a canonical party and facility. “Existing overdraft” must resolve to a booked product and limit. “Can keep” might mean contractual entitlement, policy eligibility, operational continuation or a recommendation to an approver. “Annual review” has a case state and effective date. “Open” could refer to a workflow status rather than a credit conclusion.

A safe plan therefore needs five dimensions before similarity enters the picture:

Dimension Compiler question Source of truth Unsafe inference
Principal Who is acting, through which service and delegation? Identity provider, workforce role and case assignment Trust the name or role stated in the prompt
Purpose What approved task is being performed? Workflow type, case state and purpose code Treat any work-related question as sufficient purpose
Resource scope Which customer, facility, entity and document classes are permitted? Relationship graph, entitlements and information barriers Search broadly, then remove forbidden results
Applicability Which jurisdiction, product, segment and policy versions govern? Policy registry and effective-date rules Let semantic similarity choose a policy
Evidence need What propositions must be supported, contradicted or marked unknown? Decision schema and evidence contracts Fill the token window with generally relevant text

The dimensions are related but not interchangeable. Maya may have permission to read the current credit file yet lack permission to see a restricted investigation note. A policy may be globally visible but inapplicable to the Indian entity. A customer email may be authorised and current but insufficient to prove that a waiver was approved. A relevant source can fail three separate admission tests.

The separation follows established access-control ideas. NIST SP 800-162 describes attribute-based access decisions using attributes of the subject, object, action and environment. NIST SP 800-207 frames zero trust around explicit, resource-focused decisions rather than implicit trust based on network position. Neither publication specifies a RAG architecture. They do establish a useful principle: the decision should be explicit and evaluated against the requested resource.

Relevance is not a permission, and permission is not evidence sufficiency. The compiler keeps those judgements in different stages so that each can fail visibly.

Compile a plan, not a better prompt

Prompt construction is a presentation concern. Compilation is a control concern. A prompt can say “only use policies that apply to India,” but that instruction is interpreted by the same probabilistic component that is expected to read the retrieved material. A compiler resolves the constraint before a search service receives the query. If the constraint cannot be resolved, the plan does not execute.

The compiler input is a typed envelope rather than an unstructured chat message. It carries identifiers that were resolved through approved interaction flows. The natural-language question remains available for query analysis, but it cannot override the envelope.

Figure 2Natural-language question to plan join and contradiction checksCausal and control schematic
Natural-language question to plan join and contradiction checks11 declared states connected by 12 authored relations. The figure supports the section Compile a plan, not a better prompt. L0L1L2L3L4
No
Yes
01
Natural-language question
02
Intent and entity candidates
03
Workflow envelope
04
Request canonicaliser
05
Ambiguity within threshold?
06
Ask a bounded clarification
07
Canonical decision request
08
Authorisation compilation
09
Applicability compilation
10
Evidence-contract compilation
11
Plan join and contradiction checks
Reading. The authored topology makes 12 declared relations across 11 states inspectable. Read it as the control structure for “Compile a plan, not a better prompt”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

A minimal request object includes the authenticated actor, the agent service identity, the customer and case identifiers, the requested action, the purpose, the decision time and the legal entity. Delegation matters because the human, the agent and the downstream retrieval service are distinct principals. A user’s broad application access should not become a reusable bearer capability for an autonomous process.

RFC 8693 defines OAuth token exchange patterns that can represent delegation and impersonation semantics. RFC 9396 adds structured authorization_details for fine-grained requests. These standards do not solve bank entitlement design. They do show how an application can carry a bounded, machine-readable request rather than rely on a coarse scope string.

The compiled plan might contain the following conceptual fields:

Plan field Example for atlas Control property
decision_type credit.review.continuation Selects an approved decision schema
principal_chain Maya → review agent → retrieval service Preserves delegation and service identity
purpose annual_credit_review Limits secondary use
resource_roots customer, facility and case identifiers Prevents entity drift
policy_scope entity IN-01, SME overdraft, effective 2026-07-22 Makes applicability executable
deny_labels investigation, whistleblowing, unrelated-party data Creates non-negotiable exclusions
required_propositions limit, review status, waiver, continuation rule Defines evidence coverage
retrieval_budget candidates and tokens by evidence class Controls cost and context displacement
plan_expiry short duration or case-state change Prevents stale replay
policy_hash signed bundle digest Supports reproduction and invalidation

The model may propose a search expression; it may not create or relax the authorisation predicate. Query generation stays inside a compiler-owned grammar. Unrecognised fields, missing attributes and contradictory policy results fail closed.

The compiler has six passes

Compiler language is useful because the work resembles a conventional compiler more than a chatbot. The input is parsed, names are resolved, types are checked, policies are evaluated, a plan is optimised without changing semantics, and an executable artefact is emitted. Each pass has an error class and an owner.

Figure 31 Parse to unsigned or expired dependenciesCausal and control schematic
1 Parse to unsigned or expired dependencies12 declared states connected by 9 authored relations. The figure supports the section The compiler has six passes. L0L1L2 01
1 Parse
02
2 Resolve
03
3 Type-check
04
4 Authorise
05
5 Plan
06
6 Attest
07
Malformed request
08
Ambiguous entity
09
Unsupported decision
10
Denied resource or purpose
11
Unsatisfied evidence contract
12
Unsigned or expired dependencies
Reading. The authored topology makes 9 declared relations across 12 states inspectable. Read it as the control structure for “The compiler has six passes”, not as measured performance. Dashed paths mark hypotheses, uncertainty or non-authoritative return paths. Schematic derived from the paper's authored topology; no measured quantities.

The parse pass converts the interaction into the decision-request schema. It rejects requests that combine incompatible actions, such as asking for both an eligibility assessment and an account restriction in one free-form operation. The resolve pass maps names to stable identifiers. Two customers with similar names remain two candidates until the user selects through an authorised interface.

The type-check pass determines whether the requested decision exists in the catalogue and whether the workflow permits it. “Summarise the case” and “recommend continuation” require different evidence. The authorisation pass queries the policy decision point. The plan pass chooses indices, retrievers, query forms, candidate budgets and validation rules. The attestation pass binds the output to policy and index versions.

Pass Must be deterministic? Permitted model assistance Release evidence
Parse Output validation must be deterministic Classify intent and extract candidates Schema conformance and ambiguity tests
Resolve Canonical selection must be confirmed Generate aliases or candidate queries Entity-resolution precision by risk tier
Type-check Yes Explain why a supported type may fit Negative tests for unsupported combinations
Authorise Yes None in the decision path Policy unit tests and deny-dominance tests
Plan Constraints yes; ranking choice may be learned Rewrite searches inside allowed scope Semantic-equivalence and budget tests
Attest Yes None Signature, expiry and dependency verification

This design does not require every component to be handwritten. A model can identify likely entities, classify a request and generate alternate lexical queries. Its proposals are data submitted to typed validators. Probabilistic interpretation may enrich the plan; deterministic controls decide whether the plan exists.

Policy-filtered means before candidate generation

Many systems perform “security trimming” after search. They fetch a large candidate set, remove documents whose access-control list does not include the caller, and pass the remainder to a reranker. This can be acceptable only if the retrieval service itself is authorised to see the broader set, side channels are controlled, and recall after filtering remains adequate. It is a dangerous default for mixed-sensitivity bank corpora.

Pre-filtering constrains candidate generation. The lexical query and vector search operate only over items satisfying the compiled predicate. Azure AI Search now documents several document-level access-control approaches, including permission metadata and security filters. Amazon Bedrock Knowledge Bases documents metadata filtering in retrieval configuration. These product capabilities are building blocks, not evidence that a bank’s full policy semantics fit one filter expression.

Figure 4Compiled authorisation predicate to quarantine and incident eventCausal and control schematic
Compiled authorisation predicate to quarantine and incident event9 declared states connected by 12 authored relations. The figure supports the section Policy-filtered means before candidate generation. L0L1L2L3L4
Mismatch
01
Compiled authorisation predicate
02
Lexical candidate generator
03
Vector candidate generator
04
Graph traversal
05
Bounded query forms
06
Candidate union
07
Independent permission recheck
08
Evidence-class reranking
09
Quarantine and incident event
Reading. The authored topology makes 12 declared relations across 9 states inspectable. Read it as the control structure for “Policy-filtered means before candidate generation”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Pre-filtering still needs a second check. Index permissions can lag the source, a connector can mis-map a group, or a query adapter can omit part of a predicate. The independent check reads immutable entitlement metadata attached to each candidate and compares it with the plan. It should run outside the model and before content enters a prompt, trace or feature store.

Post-filtering is a defence in depth, not permission to over-retrieve. Candidate counts, timing and score distributions can leak information even when text is later removed. Logs can also capture forbidden snippets during reranking. The safest architecture minimises the unauthorised set at every boundary.

Deny dominance must survive every join

Policy engines differ in their semantics, but the retrieval architecture needs an explicit rule for conflicts. The Cedar authorisation model documents default deny and forbid policies that override permits. Open Policy Agent provides a general policy engine over structured input. Google’s Zanzibar paper describes a globally distributed relationship-based authorisation system. Each solves a different layer; none should be copied without mapping local requirements.

The important property is deny dominance across composition. Suppose Maya may read Atlas’s credit file and the annual-review agent may retrieve case documents. A separate information-barrier rule denies the agent access to investigation notes. Joining “Maya can read customer” with “agent can retrieve case” must not erase the specific deny.

Figure 5Principal and delegation attributes to permit bounded operationCausal and control schematic
Principal and delegation attributes to permit bounded operation9 declared states connected by 9 authored relations. The figure supports the section Deny dominance must survive every join. L0L1L2L3L4
Yes
No
No
Yes
01
Principal and delegation attributes
02
Policy evaluation
03
Resource, labels and relationships
04
Action and purpose
05
Environment: time, entity, case state
06
Any applicable deny?
07
Deny with reason code
08
Required permits all present?
09
Permit bounded operation
Reading. The authored topology makes 9 declared relations across 9 states inspectable. Read it as the control structure for “Deny dominance must survive every join”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

The compiler should retain reason codes without exposing sensitive policy internals to the user. “Resource outside authorised case scope” may be safe; “document is part of an undisclosed investigation” may reveal the fact the deny protects. Operations needs a privileged diagnostic view, while the interaction receives a bounded refusal.

Test deny dominance with mutation, not examples alone. Add a new permit, change group membership, remove a label, replay an expired plan and reorder policy evaluation. The expected authorised set should remain invariant except where the change was approved. Cedar’s published design includes analyzability as an objective; the broader lesson is to treat policy equivalence and unintended widening as testable properties rather than code-review intuition.

Applicability is a separate compiler

A policy can be visible to Maya and still be the wrong authority for Atlas. Applicability depends on legal entity, jurisdiction, customer segment, product, booking location, decision type, effective date and sometimes transition provisions. Semantic retrieval is weak at this because superseded and current documents often share almost all their words.

The applicability compiler starts from a governed policy registry. Each policy version carries machine-readable scope, effective intervals, supersession links, owner, approval and source location. It emits the small set of provisions allowed to govern the decision. If two provisions conflict and no precedence rule resolves them, the plan routes to a policy owner.

Applicability field Required treatment Failure if omitted
Legal entity Match the entity making or booking the decision Group guidance can displace local rules
Jurisdiction Evaluate approved jurisdiction mapping Foreign guidance appears authoritative
Product and segment Use canonical taxonomy, not free text A retail rule enters a business case
Effective interval Evaluate against decision and event time Future or retired policy is retrieved
Supersession Follow approved replacement graph Both old and new clauses look equally relevant
Exception or waiver Require authority, scope and expiry A note is mistaken for a valid approval
Precedence Apply approved hierarchy A procedure overrides a binding policy

For Atlas, the new group manual is effective immediately for new approvals, while the local addendum keeps the prior affordability threshold for renewals until quarter end. The compiler can represent that transition because “effective for” is attached to a decision class, not only a calendar date. A recency sort cannot.

Policy selection must return a reasoned applicability receipt, not merely a top-ranked document. The receipt names the rule version, scope facts, transition rule and unresolved conflicts. The model can explain it after compilation; it cannot originate the precedence.

Figure 6Draft to suspendedCausal and control schematic
Draft to suspended8 declared states connected by 10 authored relations. The figure supports the section Applicability is a separate compiler. L0L1L2L3L4
Named authority signs
Future effective date
Immediate release
Effective condition met
Replacement applies
No replacement
Emergency control
Authorised reinstatement
01
Draft
02
Approved
03
Scheduled
04
Effective
05
Superseded
06
Retired
07
Archived
08
Suspended
Reading. The authored topology makes 10 declared relations across 8 states inspectable. Read it as the control structure for “Applicability is a separate compiler”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Evidence contracts stop context from becoming a scrapbook

Once the authorised and applicable universe is known, the compiler still has to decide what to retrieve. A generic “top 20 chunks” budget makes no distinction between a contract, a policy, a system record and an informal comment. The evidence contract defines propositions and admissible evidence classes.

For a continuation recommendation, the propositions might be: the facility exists and its approved limit; the review is open and not overdue beyond the applicable tolerance; a waiver is valid for this covenant and date; the current policy permits continuation under stated conditions; no hard stop is recorded in an authorised system. Each proposition has a source hierarchy and a contradiction rule.

Figure 7Decision schema to bounded model contextCausal and control schematic
Decision schema to bounded model context11 declared states connected by 14 authored relations. The figure supports the section Evidence contracts stop context from becoming a scrapbook. L0L1L2L3L4
No
Yes
01
Decision schema
02
Facility and limit
03
Review status
04
Waiver authority
05
Applicable continuation rule
06
Hard-stop status
07
Evidence-class queries
08
Coverage and contradiction ledger
09
Contract satisfied?
10
Unknown or human escalation
11
Bounded model context
Reading. The authored topology makes 14 declared relations across 11 states inspectable. Read it as the control structure for “Evidence contracts stop context from becoming a scrapbook”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.
Proposition Preferred evidence Secondary evidence Never sufficient alone
Approved facility limit Authoritative facility system and signed approval Versioned credit paper Email summary
Review status Workflow system event Analyst case note Natural-language mention
Waiver validity Approved waiver object with authority and interval Signed committee record Relationship-manager note
Continuation rule Applicable policy clause Approved procedure Training slide
Hard stop Authoritative control status Signed exception Absence of a retrieved warning
Customer change Verified source or approved customer evidence Case observation Model inference from transaction narrative

The contract allows unknown. If the waiver record is missing, the system should not let a similar historical waiver fill the gap. It can retrieve the historical item as comparison material only if authorised and clearly typed as non-case evidence. A context packet is complete when it exposes the decision’s evidential state, not when it fills the token budget.

The approach is consistent with the provenance vocabulary in the W3C PROV-O Recommendation, which distinguishes entities, activities and agents and includes generation and invalidation times. PROV-O does not prescribe a bank evidence contract. It supplies a useful interoperable model for saying where an artefact came from and how it was produced.

Compile queries per evidence class

One search expression rarely serves all evidence. Exact identifiers and controlled terms suit lexical retrieval. Similar wording can help find policy commentary. Graph traversal can resolve customer–facility–case relationships. Structured lookups should fetch balances, limits and status. The compiler chooses a retrieval route for each proposition rather than forcing every source into a vector index.

For Atlas, the facility-limit query is a structured read keyed by the facility identifier. The waiver query combines that identifier, covenant code and valid time. The policy query uses an exact product and entity filter, then lexical and semantic retrieval within the admitted policy versions. A graph traversal checks that the facility belongs to the resolved customer. None of these routes receives a request to “search everything about Atlas.”

Figure 8Proposition to proposition ledgerCausal and control schematic
Proposition to proposition ledger8 declared states connected by 10 authored relations. The figure supports the section Compile queries per evidence class. L0L1L2L3L4
Record state
Clause or definition
Relationship
Signed artefact
01
Proposition
02
Evidence class
03
Typed system API
04
Lexical + vector retrieval
05
Authorised graph traversal
06
Object store by stable identifier
07
Evidence normaliser
08
Proposition ledger
Reading. The authored topology makes 10 declared relations across 8 states inspectable. Read it as the control structure for “Compile queries per evidence class”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Query rewriting can remain useful. The model may expand “keep the overdraft” into approved terms such as continuation, temporary extension or review grace period. The grammar rejects fields outside the policy scope. It does not allow the model to add a customer, remove an entity filter or request a restricted source. Generated queries are logged as plan data, not trusted as instructions.

Indexes need control metadata, not decorative tags

Policy filtering succeeds only if ingestion preserves the attributes used at query time. A chunk with text and a source URL is not enough. It needs a stable resource identifier, source version, access-control representation, legal entity, policy scope, validity, document class, evidence class, lineage and deletion state. Chunk-level metadata must inherit conservatively from the source.

If one document contains sections with different access conditions, the safest choice may be to split it into separately governed resources before chunking. If a chunk joins two pages with different classifications, it inherits the stricter classification. If a parser cannot identify a document version or permission, the content goes to quarantine rather than a default-public index.

Ingestion invariant Verification Failure disposition
Every chunk maps to one stable source version Referential-integrity check Quarantine
Access labels are present and understood Schema and allow-list validation Deny indexing
Permission inheritance is conservative Source-to-chunk comparison Apply stricter label and investigate
Effective and supersession data are resolvable Interval and graph validation Exclude from live-decision index
Deletion propagates to all projections Tombstone reconciliation Block affected index partition
Embedding maps to exact normalised content Content and model-version hashes Recompute or exclude

The Azure AI Search security guidance explicitly treats document-level controls as a way to restrict retrieved documents by identity. The product documentation also reveals a practical dependency: permission information has to enter and remain aligned with the index. A compiler cannot compensate for missing or stale entitlement metadata.

Unknown classification is denied classification. “Internal” is not a useful fallback when the missing attribute might have been a customer information barrier or an investigation restriction.

The evidence manifest is the interface to the model

The compiler output is not a bag of chunks. It is a context package with typed sections and a manifest. The manifest records the plan ID, principal chain, purpose, policy hash, query routes, evidence identifiers, admission reasons, retrieval times, source versions, contradictions and omissions. Content is clearly marked as data, never as a new instruction channel.

Figure 9Authorised user to language modelInteraction sequence
Authorised user to language model6 declared states connected by 8 authored relations. The figure supports the section The evidence manifest is the interface to the model. t
Authorised user
Context compiler
Policy decision point
Retrieval services
Evidence validator
Language model
01
Typed decision request
02
Principal, action, resource, purpose, environment
03
Bounded permits and denies
04
Signed retrieval plan
05
Candidates with source metadata
06
Coverage, conflicts and rejected items
07
Evidence manifest plus admitted excerpts
08
Draft answer with proposition citations and unknowns
Reading. The authored topology makes 8 declared relations across 6 states inspectable. Read it as the control structure for “The evidence manifest is the interface to the model”, not as measured performance. Dashed paths mark hypotheses, uncertainty or non-authoritative return paths. Schematic derived from the paper's authored topology; no measured quantities.

The manifest lets the model cite propositions rather than simply repeat document titles. It can say the limit is supported by record F-102 as retrieved at a specified time, while waiver validity is unresolved because no approved waiver object was admitted. The interface should prevent the answer from citing rejected candidates or references seen during query generation.

Indirect prompt injection remains a residual risk because retrieved content can contain adversarial instructions. The OWASP prompt-injection prevention guidance describes attacks delivered through external content and recommends layered controls. The compiler contributes containment: retrieved text cannot alter permissions, acquire new tools or expand the plan. It does not make the model immune to manipulation.

Documents supply evidence, never authority. Instructions found in a policy PDF, email, web page or customer upload cannot change the compiler plan, system prompt, tool set, recipient or output destination. Any request to perform an action is treated as quoted content unless it arrives through the authorised control channel.

Do not put secrets in the plan

Auditability can create a second data leak if the plan or trace copies sensitive attributes broadly. The signed plan needs stable identifiers and policy results, not every group membership, customer detail or deny-rule body. A privileged service can resolve references during an investigation. Routine application logs should contain reason codes and hashes sufficient for correlation.

Traces also require purpose and retention controls. A reranker trace that stores candidate snippets may retain customer data outside the source system. A prompt-observability platform may become a shadow evidence repository. The compiler should specify which fields may be logged, tokenised or redacted, and the enforcement point should apply that policy before export.

The NIST AI Risk Management Framework and its Generative AI Profile are voluntary cross-sector references rather than bank-specific rules. Their emphasis on mapping, measuring and managing risks across the lifecycle supports a simple architectural conclusion: evidence handling, privacy and monitoring must be designed into the system, not appended to the answer screen.

Failure paths are part of the product

A compiler that works only on complete metadata will fail in a real bank. The design must make incomplete identity, stale entitlements, ambiguous entities, policy conflicts, index lag and missing evidence predictable. “Try a broader search” is not an acceptable generic recovery.

Failure Safe system response Human option Evidence retained
Principal or delegation cannot be verified Deny execution Reauthenticate or use approved manual process Authentication and reason code, no content
Customer maps to multiple entities Stop before retrieval Select through authorised case UI Candidate IDs without unauthorised attributes
Policy scope conflicts Do not choose by rank Route to policy owner Conflicting rule versions and facts
Index permission version lags source Exclude partition or use source lookup Retrieve manually in source system Version mismatch and affected resources
Required evidence is absent Return unknown, not a guessed answer Request evidence or escalate Queries, admitted sources and gap
Post-check finds forbidden candidate Drop candidate and open security event Security investigates connector or policy Candidate ID, plan, policy and index versions
Plan expires during long task Stop further reads and recompile Confirm continuation Prior plan and new decision receipt

A denied or incomplete plan is a valid product outcome. Its usability matters. The response should state what cannot be established, which safe step can resolve it and which decision remains reserved for a person. It should not reveal the existence of a protected record.

Figure 10Compilation or validation error to approved manual fallbackCausal and control schematic
Compilation or validation error to approved manual fallback8 declared states connected by 7 authored relations. The figure supports the section Failure paths are part of the product. L0L1L2
Identity or authority
Ambiguity
Policy conflict
Evidence gap
Index inconsistency
Dependency outage
01
Compilation or validation error
02
Classify failure
03
Deny and reauthenticate
04
Bounded clarification
05
Policy-owner queue
06
Return unknown with gap
07
Stop partition and security event
08
Approved manual fallback
Reading. The authored topology makes 7 declared relations across 8 states inspectable. Read it as the control structure for “Failure paths are part of the product”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Test the authorised set before answer quality

Most retrieval evaluation starts with relevance judgements. A context compiler needs an earlier suite: given a principal, purpose, resource and environment, did it produce exactly the permitted candidate universe? False admission and false denial have different costs, and a single average hides both.

Construct policy fixtures with known resource sets. Include nested groups, revoked delegation, cross-entity assignments, temporary access, information barriers, policy transitions and deletion. Run each fixture against the compiler and every retrieval adapter. Mutation tests should remove one predicate, alter a label and replay an old plan. Any widening must fail the release.

The compilation suite also tests evidence contracts. For each decision type, hide one authoritative source, insert a contradictory source, provide only secondary evidence and include an inapplicable but highly similar policy. The expected output is a gap, conflict or exclusion: not a lower confidence score attached to a complete-sounding answer.

Test family Primary measure Release condition Why answer metrics cannot replace it
Authorisation soundness False-admission rate by sensitivity Zero in deterministic fixture set A fluent answer can conceal a forbidden read
Authorisation completeness False-denial rate by role and case Within approved operational bound Low recall may block legitimate work
Applicability Correct governing version and precedence All golden transitions pass Similar text is not governing authority
Evidence coverage Required propositions with admissible support No unsupported “known” proposition Citation presence does not prove sufficiency
Contradiction handling Conflicts preserved and routed No silent collapse in critical propositions Majority voting can erase an authoritative exception
Replay resistance Expired or changed plan rejected All stale-plan tests pass An answer can be accurate under an obsolete policy
Non-interference Forbidden content cannot change tools or scope All injection cases contained Content filters do not enforce permissions

Only after these pass should the team measure retrieval relevance and answer usefulness. The first quality question is “was this evidence allowed and applicable?” The second is “was it useful?”

Walk the atlas case through the compiler

The Northstar case becomes useful when every abstract control changes a concrete system action. Maya opens the approved annual-review workspace. The workspace already knows the customer, case and legal entity. The assistant receives those identifiers through the workflow envelope; it does not recover them from conversation history. Maya’s identity token is exchanged for a short-lived agent capability restricted to reading evidence for this case.

The parser classifies the request as a continuation assessment, not a credit approval. Entity resolution verifies that the overdraft belongs to Atlas and that the case concerns the same booking entity. The policy decision point permits the agent to read the facility record, approved credit paper, annual-review state and applicable policy. It denies the restricted note. The deny is not disclosed in the interaction.

The applicability compiler selects the group manual and local transition addendum. It excludes a future-dated local policy and a retired training guide. The evidence compiler opens five proposition slots. It uses a structured lookup for the facility limit and case status, an object lookup for the waiver, a filtered policy search for continuation rules, and an authorised status service for hard stops.

The waiver object presents the decisive failure. Its authority is valid, but its valid_to time passed two days earlier. A relationship-manager note says a renewal was requested. That note is permitted context, yet it does not meet the evidence contract for an approved waiver. The packet marks the waiver proposition unresolved. The draft answer can explain that continued availability cannot be established from the admitted evidence and can identify the approved route for a credit officer to resolve it.

Figure 11Maya to language modelInteraction sequence
Maya to language model6 declared states connected by 9 authored relations. The figure supports the section Walk the atlas case through the compiler. t
Maya
Review workspace
Context compiler
Authorisation service
Evidence sources
Language model
01
Ask continuation question
02
Case-bound request and identity
03
Compile principal, purpose and resources
04
Permit bounded reads; apply specific deny
05
Execute proposition-specific plan
06
Limit, status, expired waiver, policy
07
Manifest: four supported, one unresolved
08
Draft assessment; no approval or execution
09
Evidence view and officer escalation
Reading. The authored topology makes 9 declared relations across 6 states inspectable. Read it as the control structure for “Walk the atlas case through the compiler”, not as measured performance. Dashed paths mark hypotheses, uncertainty or non-authoritative return paths. Schematic derived from the paper's authored topology; no measured quantities.

The outcome is less convenient than an unconditional answer. It is also more useful. Maya knows exactly which issue blocks the assessment. The bank has not converted a pending request into an approval, and no protected content has entered the model context. If an authorised officer approves a new waiver, a source event invalidates the plan and the question can be recompiled against the new state.

Evaluation, assurance and counterevidence

Cache plans only under their dependencies

Compilation can add latency, so teams will cache plans and policy decisions. The cache key must include every dependency that can change the authorised set: principal and delegation version, purpose, resource root, policy bundle, relevant group memberships, information-barrier state, case state and effective time. A query string alone is never a safe key.

A plan carries a short expiry and a dependency list. Source events can revoke it earlier. If the user changes role, the case transfers, a customer receives a new classification, a policy becomes effective or a document ACL changes, the relevant plans are invalidated. The next tool call cannot rely on a permission check performed at the beginning of a long conversation.

Figure 12Compiled plan to recompileCausal and control schematic
Compiled plan to recompile9 declared states connected by 8 authored relations. The figure supports the section Cache plans only under their dependencies. L0L1L2L3
Valid dependency versions
Expired or invalidated
01
Compiled plan
02
Dependency key
03
Short-lived plan cache
04
Identity and group events
05
Invalidation bus
06
Resource and ACL events
07
Policy and case-state events
08
Reuse bounded plan
09
Recompile
Reading. The authored topology makes 8 declared relations across 9 states inspectable. Read it as the control structure for “Cache plans only under their dependencies”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

HTTP caching concepts are informative but insufficient. RFC 9111 distinguishes freshness from validation and makes origin controls explicit. An authorisation plan has more dependencies than a representation’s age. Its apparent freshness does not prove that the entitlement graph or case state is unchanged.

Treat cache failure asymmetrically. A cache miss costs time. A stale permit can expose data. During an invalidation outage, high-sensitivity reads should re-evaluate at the source of authority or stop. A low-risk knowledge search may have a documented bounded fallback. Availability policy must never be an accidental “permit on error.”

Measure the compiler as a control plane

Answer acceptance is a weak operational metric. Users may accept a convenient answer drawn from the wrong source. Compiler telemetry should show whether plans are sound, complete, timely and explainable. It should also reveal where source governance, not model quality, limits performance.

Define measures at each pass. Parsing needs schema rejection and clarification rates. Resolution needs ambiguous-entity and wrong-entity rates from reviewed samples. Authorisation needs permit, deny and error counts by policy version, plus false-admission findings from test and audit. Applicability needs conflict and unresolved-precedence rates. Evidence planning needs proposition coverage, source-class substitution and contradiction preservation. Execution needs candidate leakage and post-check mismatch counts.

Metric Definition Interpretation Anti-gaming guard
Plan compilation success Valid plans / eligible requests Usability of schemas and dependencies Exclude unsupported requests explicitly; do not relabel denials as errors
Unsafe admission Forbidden candidates admitted before model boundary Security-control failure Count every candidate, not only cited ones
Permission mismatch Candidates rejected by independent recheck Index or adapter inconsistency Alert on any sensitive mismatch
Governing-version accuracy Correct policy version / sampled decisions Applicability quality Sample policy transitions disproportionately
Evidence-contract coverage Supported required propositions / required propositions Readiness for a decision Keep unknown separate from failure
Unsupported proposition rate Answered propositions lacking admitted evidence Grounding failure Evaluate structured claims, not citation count
Denial usefulness Denials with an approved resolution route / denials Operational design quality Do not reveal protected-resource existence
Compilation latency Pass-level p50, p95 and p99 User experience and dependency health Report cache-hit and cache-miss paths separately
Plan invalidation lag Event time to unusable plan Stale-authority exposure Test with synthetic revocation events
Manual fallback rate Eligible requests completed outside compiler Adoption and resilience signal Review whether fallback bypassed equivalent controls

No universal threshold belongs in an architecture paper. Northstar could set a release rule of zero false admissions in a deterministic security fixture, full passage of policy-transition cases and no unsupported critical proposition in the golden set. Those are proposed test conditions, not measured achievements.

Operational dashboards should preserve denominators. A fall in compiler success might be caused by a policy release that introduced unresolved scopes. A rise in evidence gaps might indicate a connector outage, not a weaker retriever. Measure the pass that failed; do not turn every defect into a model-accuracy problem.

Reconcile access at source, index and context boundaries

Three enforcement layers can drift. The source system owns the canonical entitlement. The index carries a projection used for candidate generation. The context validator checks each result before the model boundary. The bank needs continuous reconciliation among them.

Take samples and compare source permissions with index metadata. Include recently revoked users, newly classified documents, nested groups and resources moved between cases. Recompute an expected authorised set from policy fixtures and compare it with results from each index. Reconciliation should cover deletion and retention as well as positive access.

Figure 13Source permissions to re-run security fixturesCausal and control schematic
Source permissions to re-run security fixtures10 declared states connected by 10 authored relations. The figure supports the section Reconcile access at source, index and context boundaries. L0L1L2L3L4
No
Yes, isolated
Yes, systemic
01
Source permissions
02
Reconciliation job
03
Index permission projection
04
Context admission receipts
05
Mismatch?
06
Record coverage and version
07
Quarantine resource
08
Stop affected partition
09
Repair and reindex
10
Re-run security fixtures
Reading. The authored topology makes 10 declared relations across 10 states inspectable. Read it as the control structure for “Reconcile access at source, index and context boundaries”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

The reconciliation service must not become a broad content reader. It can compare stable resource IDs, permission hashes, labels and versions. When a mismatch requires inspection, a privileged workflow opens only the affected resource. Results feed connector ownership and incident management.

Banking data governance provides a relevant operational frame. The Basel Committee’s BCBS 239 principles concern effective risk-data aggregation and reporting, not generative retrieval. The Committee’s January 2026 implementation newsletter reiterates the importance of accurate, comprehensive and timely data and explicitly says it creates no new supervisory expectations. The useful inference is limited: an agentic context layer should not weaken lineage, ownership or reconciliation expected of important bank data.

Separate compiler policy from retrieval tuning

Retrieval engineers will tune candidate counts, fusion weights, rerankers and chunking. Security and policy owners will change entitlement and applicability rules. If both are packed into one prompt or configuration file, a relevance experiment can widen access accidentally.

Maintain separate artefacts. The policy bundle determines the admissible universe and evidence classes. The retrieval profile chooses how to rank inside that universe. The model profile determines how admitted evidence is presented and summarised. A release manifest binds their versions but preserves independent ownership and rollback.

Artefact Owns Must not own Approval
Authorisation policy Principal, action, resource, purpose and environment rules Search weights or wording Security and data owner
Applicability policy Scope, effective time, supersession and precedence User entitlements Policy owner and legal/compliance interpretation
Evidence contract Propositions and admissible evidence classes Model style Decision-process owner
Retrieval profile Query routes, candidate budgets, fusion and reranking Access widening Search engineering with evaluation evidence
Model profile Output schema, citation behaviour and abstention wording Tool or resource permissions Product and model governance
Release manifest Compatible, tested version set Hidden overrides Joint change authority

The separation makes experiments safer. A new embedding model can run against the same authorised candidate fixtures. A revised policy can be tested with the existing retrieval stack. A prompt change cannot alter resource scope because the retrieval adapter accepts only a signed plan.

Optimisation must preserve policy semantics

A conventional compiler optimises an intermediate representation without changing program meaning. A context compiler needs the same constraint. It may combine equivalent filters, route a structured lookup directly, reuse a valid policy decision or lower candidate counts after coverage is achieved. It may not drop a “redundant” deny, approximate an effective-time boundary or trade authorisation completeness for latency.

Semantic-equivalence tests compare the resource set before and after an optimisation. Where the universe is finite, enumerate fixtures. Where relationship graphs are large, use property-based tests and targeted proofs for rule classes. Include empty and unknown attributes because optimisers often mishandle three-valued logic.

An attractive optimisation is to retrieve broadly from a shared vector index and rely on a fast post-filter. Reject it if it changes which services process forbidden candidates or if score and timing data can escape. Another is to cache group expansion. Accept it only with bounded freshness, event invalidation and a deny-on-uncertain policy appropriate to the resource class.

Performance work can change execution plans, never the authorised meaning of a request. That line belongs in design review, automated tests and incident criteria.

Plan for policy distribution failure

The policy decision point and registry become critical dependencies. A single central service can add latency and create an outage domain; distributing policy can introduce version skew. The choice needs an explicit consistency model.

One pattern distributes signed policy bundles to local evaluators. The compiler checks the bundle version against a minimum accepted version and includes the digest in every plan. Emergency denies travel through a priority channel and invalidate affected plans. New permits can tolerate slower propagation than revocations because delay reduces availability rather than widening access.

Figure 14Approved policy origin to receipts include bundle digestCausal and control schematic
Approved policy origin to receipts include bundle digest10 declared states connected by 13 authored relations. The figure supports the section Plan for policy distribution failure. L0L1L2L3 01
Approved policy origin
02
Signed bundle registry
03
Local evaluator A
04
Local evaluator B
05
Local evaluator C
06
Emergency deny channel
07
Compiler instance
08
Compiler instance
09
Compiler instance
10
Receipts include bundle digest
Reading. The authored topology makes 13 declared relations across 10 states inspectable. Read it as the control structure for “Plan for policy distribution failure”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

If an evaluator lacks a required bundle, the compiler should not silently use the newest bundle it happens to have. It can deny, route to an approved source-system workflow or use a specifically authorised degraded mode. The degraded mode must state which decision types and sensitivity classes it supports.

Counterevidence: where a compiler adds too much machinery

Not every retrieval application needs this architecture. A public website assistant over a single, openly licensed and non-versioned corpus may need source quality and injection controls but no workforce entitlement compiler. A personal search tool where one user owns every document may use operating-system permissions and a simpler freshness filter. A deterministic database view can be superior to a compiler for one stable query.

The architecture also cannot repair weak source governance. If policies have no scope metadata, entitlements exist only in managers’ memories, and customer identifiers conflict across systems, compilation will surface many denials and gaps. That friction is evidence of unresolved governance, but it can make the product unusable until the source work is funded.

Formal policy languages do not remove interpretation. A perfectly evaluated rule can encode the wrong legal or business decision. Relationship-based models can become difficult to explain when nested deeply. Attribute-based models depend on accurate attributes. Central policy services can concentrate operational risk. Fine-grained predicates can damage retrieval recall when metadata is incomplete.

There is also a human-factors limit. A long receipt can overwhelm an analyst and create ritual approval. The interface should show the few facts that change the decision, expose conflicts, and allow drill-down. It should not ask the user to inspect every policy evaluation.

Use a context compiler where access, applicability and evidence vary per request and a wrong admission matters. Do not turn the pattern into a ceremonial gateway for low-risk public search.

A release programme in four increments

The bank should not begin by compiling every policy and indexing every source. Start with one decision type whose authoritative evidence is identifiable and whose human authority is unchanged. The Atlas continuation assessment is suitable because it can produce a draft and evidence gaps without executing a credit decision.

Increment one compiles identity, case scope and source-system reads. It uses no generative query rewriting. Increment two adds the applicability registry for a bounded policy family. Increment three introduces lexical and vector retrieval inside the admitted policy set, with independent candidate checks. Increment four adds model-assisted entity and query proposals, still contained by the grammar.

Increment Capability Entry evidence Exit evidence
1. Case-bound records Principal chain, resource roots, structured evidence Canonical IDs and entitlements Security fixtures, replay and manual-fallback tests
2. Policy applicability Scope, effective time, supersession and receipt Approved policy registry Golden transition cases and conflict routing
3. Filtered hybrid retrieval Pre-filtered candidate generation and post-check Complete control metadata Leakage tests, retrieval evaluation and reconciliation
4. Model-assisted planning Bounded intent, entity and query proposals Stable compiler grammar Adversarial, ambiguity and drift evaluation

Run shadow compilation against real requests without exposing results to the model. Compare the compiler’s admitted set with authorised reviewers’ source choices. Investigate differences by cause: policy, entity, entitlement, evidence class or search. Do not use reviewer behaviour as unquestioned ground truth; a reviewer may have relied on an inapplicable document.

Promotion needs named acceptance evidence. Security owns false-admission criteria. Policy owners own applicability fixtures. The decision-process owner owns evidence contracts. Search engineering owns relevance inside the permitted set. Operations owns latency, recovery and reconciliation. Internal audit or independent assurance should be able to reperform a sample from the plan receipt.

Implementation and operating detail

Operating decisions to record before production

Architecture diagrams do not resolve trade-offs. The programme should record decisions in a concise ledger and revisit them when scope changes.

Decision Conservative default Evidence required to change it
Pre-filter versus post-filter Pre-filter all sensitive corpora; recheck every candidate Threat analysis showing equivalent isolation and no side-channel exposure
Unknown entitlement attribute Deny and quarantine Approved, resource-specific fallback with monitoring
Policy conflict Route to named owner Approved precedence encoded and tested
Missing critical evidence Return unknown Alternative evidence class approved in contract
Plan lifetime Short and event-invalidated Measured dependency stability and accepted exposure analysis
Model-generated filters Prohibited for security and applicability None; allow only bounded query terms inside compiled predicates
Trace content Identifiers, reason codes and hashes by default Purpose, access and retention approval for excerpts
Service outage Source-system manual route Approved degraded mode with equal control boundary

These are not universal prescriptions. They are a starting risk posture for the composite use case. A bank may choose differently after documenting data sensitivity, service criticality and control evidence.

Give the plan an intermediate representation

Directly translating one interaction into one vendor query makes the control hard to inspect and hard to move. A better design creates a vendor-neutral intermediate representation between policy evaluation and retrieval. The representation expresses sets and constraints, not search syntax. Adapters lower it into an SQL predicate, a graph traversal, a lexical filter or a vector-store request.

The intermediate representation needs a small type system. CustomerId must not be interchangeable with FacilityId. A DecisionTime differs from an ingestion timestamp. PolicyVersion is not a document title. Evidence classes form a controlled enumeration. Labels and jurisdictions come from governed namespaces. The compiler rejects a query adapter that cannot preserve a required predicate.

For example, an applicability expression could require entity IN-01, product class SME_OD, decision class REVIEW_CONTINUATION, and an effective interval containing the decision time. An authorisation expression could require an allowed relationship between the principal and case, the purpose ANNUAL_REVIEW, and the absence of any applicable deny label. The intermediate form retains these as separate clauses even if a search product combines them into one filter.

Intermediate type Valid operations Invalid coercion Why the distinction matters
Stable resource ID Equality, membership, relationship traversal Fuzzy text match Prevents names from silently switching customers
Classification label Ordered comparison or explicit policy relation Treating missing as public Preserves conservative inheritance
Valid-time interval Contains, overlaps, precedes Sorting by upload date Selects governing versions rather than recent files
Evidence class Approved substitution relation Converting any citation into proof Keeps commentary distinct from authority
Purpose code Exact approved value and hierarchy Inferring purpose from question alone Limits secondary use of accessible data
Principal chain Verified delegation edges Copying the human’s full token Keeps service authority bounded
Budget Candidate and token ceilings per class Moving budget between classes silently Stops abundant commentary displacing a required record

Adapter conformance is a release gate. Feed the adapter synthetic plans with nested booleans, empty groups, denied resources, interval boundaries and non-ASCII identifiers. Capture the actual candidate set. Compare it with a reference evaluator. An adapter that approximates a filter for performance must declare that loss and must not be used for sensitive classes.

The representation should also make non-monotonic decisions visible. Adding a deny label can shrink the authorised set; adding a new approved evidence source can expand available support without expanding customer scope. The compiler records which change caused a new plan. This helps assurance teams distinguish a retrieval improvement from a policy widening.

Multi-turn conversation must recompile material changes

Conversation creates pressure to reuse context. Maya may follow the Atlas question with “What about its parent?” or “Compare it with last year.” A chat system may treat those as harmless continuations. They change the resource root, evidence contract or decision time.

The session should store a reference to the prior plan, not a blanket permission for all subsequent turns. The request canonicaliser compares each follow-up with the active envelope. A wording clarification can reuse the plan if dependencies and purpose remain unchanged. A new customer, facility, time basis, action or output recipient requires recompilation. A case-state event can force recompilation even when Maya repeats the same words.

Figure 15Noplan to deniedCausal and control schematic
Noplan to denied5 declared states connected by 8 authored relations. The figure supports the section Multi-turn conversation must recompile material changes. L0L1L2L3
Compile bounded request
Pure wording clarification
Resource, purpose, action or time changes
Dependency invalidation event
Plan lifetime ends
New plan permitted
New scope not permitted
User continues task
01
NoPlan
02
ActivePlan
03
Recompile
04
Expired
05
Denied
Reading. The authored topology makes 8 declared relations across 5 states inspectable. Read it as the control structure for “Multi-turn conversation must recompile material changes”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Context from the prior turn follows the same rule as new retrieval. If an earlier excerpt is outside the new plan, it must not remain in the model’s effective context. This can require a fresh model call rather than continuing a provider-managed thread whose hidden state cannot be selectively removed. The architecture team needs to know whether its model API actually permits deterministic context reconstruction.

Agent-to-agent delegation requires another compile boundary. A review agent may ask a policy-specialist agent to explain a clause. The child receives a capability limited to the relevant policy resources and purpose; it does not inherit the full customer packet. Its result returns as untrusted analytical material with lineage. If a child needs a new source class, the parent cannot grant it through prose.

Conversation continuity is a usability feature, not an authorisation primitive. Plans, not chat sessions, carry authority.

Control the relationship between query and identity

Search logs can reveal customer interests even when results are protected. An analyst’s question may contain a customer name, account number, medical detail or investigation clue. Query rewriting can replicate those details across model providers and telemetry services. The compiler must minimise query content as well as retrieved content.

Structured identifiers should stay in trusted retrieval calls. A policy-search query may need product and decision terms but not the customer’s name. A general language model can propose synonyms from a redacted decision description. Where an external model service is permitted, the service receives only the fields approved for that processing arrangement. The plan records which transformation saw which data class.

Tokenisation is not always enough. Stable tokens can be linkable across traces. Hashing a small identifier space can be reversible by enumeration. Redaction can remove the term that made lexical retrieval effective. These trade-offs belong in the threat model and evaluation, not in a generic “PII removed” flag.

Query observability should separate operational facts from content. Teams can measure latency, candidate count, adapter errors and policy version without storing raw questions. Samples for relevance review need an approved purpose, access, retention and de-identification process. A user-interface control should not promise that “chats are not stored” if downstream search traces retain them.

Threat-model the compiler itself

Moving controls out of the prompt reduces one class of risk but creates a valuable control service. An attacker may target entity resolution, policy inputs, source metadata, plan signatures, adapter logic, caches or invalidation events. An insider may construct an allowed purpose around an illegitimate objective. A compromised connector may label restricted documents as public.

Figure 16Goal: place forbidden or inapplicable evidence in context to prevent, detect, contain, recoverCausal and control schematic
Goal: place forbidden or inapplicable evidence in context to prevent, detect, contain, recover9 declared states connected by 14 authored relations. The figure supports the section Threat-model the compiler itself. L0L1L2 01
Goal: place forbidden or inapplicable evidence in context
02
Forge principal or delegation
03
Manipulate purpose or resource resolution
04
Poison policy or entitlement attributes
05
Exploit query adapter or filter parsing
06
Replay stale signed plan
07
Race ACL change and index update
08
Smuggle content through logs or reranker
09
Prevent, detect, contain, recover
Reading. The authored topology makes 14 declared relations across 9 states inspectable. Read it as the control structure for “Threat-model the compiler itself”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

Threat controls should map to boundaries. Strong workload identity and token audience restrict who can submit plans. Canonical workflow envelopes make purpose harder to self-assert. Policy bundles are signed and changes require named approval. Plans are signed for a specific adapter, resource root and expiry. Adapters use parameterised filters rather than concatenated query syntax. Index and source versions are reconciled. Rerankers process only admitted candidates. Egress policy prevents evidence from reaching unapproved destinations.

Detection requires canaries and invariants. Place synthetic high-sensitivity resources in test partitions with deny rules for all normal users. Any retrieval is an incident. Monitor sudden changes in authorised-set size by role, unusual clarification loops, repeated denied entity probes and plan reuse across cases. These signals need careful privacy treatment and should not become an employee-surveillance shortcut.

Recovery is not simply rotating a model key. The bank may need to invalidate plans, stop an index partition, restore an entitlement projection, identify model calls that received affected evidence and notify data owners. The evidence manifest makes that impact query possible if resource IDs and plan dependencies were retained accurately.

Design reviewer explanations from receipts

An analyst does not need a policy-engine trace. They need to know what evidence governed the draft and where judgement remains. The explanation layer should transform receipts into a stable view without inventing reasons.

For Atlas, the view can show: facility limit confirmed from the facility system; annual review open from the workflow; continuation rule selected from the group manual plus local transition addendum; waiver expired at a specific time; renewal request observed but not approved; continuation conclusion unresolved. Each statement links to an admitted source that Maya can open under her own rights.

The interface should distinguish absence from denial. If a source was searched and no approved waiver was found, say evidence was not found in the named sources at the retrieval time. If another resource was denied, do not imply it exists. If a source service was unavailable, show that the search was incomplete. These are materially different reasons for uncertainty.

Receipt state Reviewer wording Wording to avoid
Supported “Confirmed by [source/version] as of [retrieval time]” “The system knows”
Contradicted “Authorised sources disagree on [proposition]” “Most sources say”
Not found “No admissible evidence was found in [searched sources]” “No evidence exists”
Source unavailable “The packet is incomplete because [source class] was unavailable” “Likely unchanged”
Inapplicable Usually omit; privileged drill-down may show exclusion reason Presenting it as an alternative authority
Denied “The requested scope cannot be processed in this workflow” Confirming a protected document or case exists
Human judgement “Decision reserved for [authorised role]” “Please approve the agent’s decision”

Explanations should be generated from structured states and approved phrases where possible. A model can improve readability, but a validator checks that it did not convert “not found” into “does not exist” or a draft recommendation into an approval. Reviewers can correct an explanation without changing the evidence record.

Explain the evidential state, not the model’s private reasoning. A chain-of-thought transcript is neither needed for reperformance nor a substitute for source and policy receipts.

Budget context by decision value

Token ceilings are not merely cost controls. They determine which evidence competes for attention. A single long policy appendix can crowd out a facility record or contradiction. The evidence contract should allocate budgets by proposition and class before retrieval results arrive.

Reserve space for the request, decision schema, proposition ledger and governing clauses. Give authoritative case records priority. Add explanatory material only after required slots are covered. If two chunks repeat the same clause, keep the best provenance and remove redundancy. Preserve contradictory evidence even if its similarity score is lower.

The compiler can use extractive condensation for long documents, but the excerpt must remain traceable to offsets in a fixed source version. A model-authored summary is a derived artefact, not the original evidence. It carries its own lineage and should not replace decisive language when exact wording matters.

Budget tests vary document length, repeated boilerplate, number of contradictions and query complexity. The expected proposition coverage should remain stable. Measure citation displacement: how often a required item was retrieved but omitted from the final context because another class consumed the budget. This is distinct from retrieval recall.

Procurement questions that expose architectural gaps

A managed knowledge service may advertise metadata filters, hybrid search and citations. Those features do not establish control completeness. Procurement and architecture review should ask how filters interact with approximate nearest-neighbour search, how permission changes propagate, whether rejected candidates enter logs, how source versions are identified, and whether provider-managed conversations can be reconstructed without old context.

Ask whether the service supports compound predicates, nested groups, deny semantics, effective intervals, stable resource IDs and customer-managed policy decisions. Determine whether filters apply before candidate generation or only after a broad vector search. Inspect limits on filter size and list membership. Establish what happens when metadata is missing or malformed.

The portability plan should assume that product semantics differ. AWS, Azure, Google Cloud, OpenSearch and specialist vector databases expose different filters, consistency behaviours and observability. A vendor adapter can be certified for a subset of decision types. The compiler should refuse a plan whose required semantics exceed that subset.

Review question Acceptable evidence Warning sign
Where is the permission predicate enforced? Documented execution stage plus adversarial test “The prompt instructs the model not to use it”
Can denied candidates reach provider logs? Data-flow diagram and contractual/technical control Only final citations are filtered
How are ACL changes invalidated? Event path, bounded lag and reconciliation evidence Periodic full reindex with unknown lag
Can a plan be replayed after role change? Audience, expiry and dependency checks Long-lived API key represents the user
Are source versions stable? Immutable ID, content hash and temporal metadata URL and last-modified string only
Can context be rebuilt per turn? Explicit request payload and no hidden retained evidence Opaque server-side thread with broad memory
What filter semantics are unsupported? Published adapter capability matrix “All metadata filtering is supported”

This review is deliberately product-agnostic. The presence of an official feature is useful, yet the bank remains responsible for its own resource model and control objective.

Independent assurance should reperform a plan

Assurance needs a sample that can be replayed without depending on the original conversation screen. The sample package includes the canonical request, principal references, policy bundle digest, plan intermediate representation, adapter versions, admitted resource IDs, rejected-candidate reason classes, evidence ledger and output. Sensitive content remains in governed sources and is opened only by authorised reviewers.

The reviewer first checks whether the request was valid for the workflow. They then evaluate the expected authorised set from source entitlements and policy. Next they check applicability and evidence coverage. Only then do they inspect answer wording. This order prevents a good answer from excusing a bad read.

Reperformance has temporal limits. Group membership, source content and policies may have changed. The receipt needs historical references or snapshots sufficient to reconstruct the decision boundary. If the source system cannot provide past entitlement state, the bank should state that assurance limitation rather than imply complete reproducibility.

Independent testing should include cases selected for risk, not only random traffic. Sample cross-entity users, policy-release windows, recently revoked access, restricted labels, large nested groups, ambiguous customers and cases with known contradictions. Report findings by failure layer and severity. A single “RAG accuracy” score is not an assurance conclusion.

Decision-grade retrieval is demonstrable only when an authorised reviewer can reperform why each item was admitted.

Fine-tuning cannot replace compilation

A team may try to teach the model not to reveal restricted information, to prefer current policy, or to distinguish authoritative documents. Fine-tuning can improve classification and response behaviour. It cannot evaluate the caller’s live entitlements, observe a policy suspension that happened after training, or prevent an upstream retriever from sending forbidden text to the model.

Training data can encode useful terminology and decision patterns. The trained model remains one component inside the boundary. It receives an already filtered evidence packet and returns a typed draft. If it proposes a resource or tool call, the compiler evaluates a fresh plan. A model’s prior familiarity with a customer-like name or policy does not authorise its use.

Fine-tuning also complicates deletion and correction. A source removed from retrieval can remain statistically represented in weights. Do not train on customer or restricted policy content unless the processing purpose, legal basis, deletion strategy, evaluation and access controls are separately approved. Retrieval gives the institution a clearer version and revocation boundary for volatile knowledge.

Put durable language capability in the model; put changing authority, scope and evidence in the compiler.

Contract tests between compiler and adapters

The plan interface is a security boundary. Each retrieval adapter should implement a shared conformance suite before it can serve a decision type. The suite is more demanding than schema validation because two backends can parse the same expression and return different sets.

Start with a synthetic corpus whose resources exercise every policy dimension: nested groups, direct and inherited permission, deny labels, entity scope, valid-time boundaries, deleted records, missing metadata and source families. Generate plans from fixtures and compare actual results with a small reference evaluator. Include Unicode identifiers, empty lists, large group expansions and operator precedence.

Test error semantics. An unsupported operator must fail the plan, not be ignored. A truncated filter must fail. A backend timeout must not return unfiltered cached candidates. A partial result must carry an incomplete status that prevents the evidence contract from appearing satisfied. The adapter should echo a canonical predicate digest so the compiler can verify which plan it executed.

Performance tests run after semantic conformance. Optimised batch filters, local caches and index partitions are acceptable if the candidate set remains equivalent. Keep a small continuous canary suite in production to catch product or configuration changes.

Context boundaries across media

Evidence is not only text. A bank agent may retrieve a scanned signature page, a spreadsheet cell, a call transcript or a chart. The same admission rules apply, but the package needs media-specific provenance and extraction evidence.

For a table, retain row, column, unit and source sheet. For an image, retain page region and the OCR or vision extraction run. For audio, retain time offsets and transcription version. A model-generated caption or transcript is a derived assertion, not a source fact. Sensitive pixels or audio must not cross the model boundary merely because the extracted text looks harmless.

The compiler can choose a trusted local extraction tool for restricted media, then admit only the necessary derived spans. If a reviewer needs the original, the interface opens it directly from the governed source under the reviewer’s identity. The context manifest links the two.

Media conversion can also hide prompt injection or classification marks. Inspect layers, attachments, comments and hidden spreadsheet cells according to source type. No universal text sanitiser is sufficient.

Decide what happens when the compiler is wrong

The architecture reduces risk only if people can challenge it. Every evidence item needs a “wrong source,” “wrong scope,” “missing evidence” or “permission concern” feedback path. The feedback opens a governed issue; it does not instantly retrain the model or change policy.

Triage by boundary. A wrong customer mapping goes to entity-data ownership. A missing policy version goes to the registry. A forbidden item goes to security incident response. A weak query goes to search engineering. A mistaken evidence contract goes to the decision-process owner. The case remains linked to the issue and any corrected packet.

Corrections should be prospective and historical. New requests use the repaired rule or metadata. Past packets dependent on the defect can be identified through plan receipts and reviewed by materiality. Never overwrite the original packet; append a correction and outcome.

This challenge loop also supplies better evaluation cases. Promote a de-identified failure into the regression suite only after the correct expected behaviour is adjudicated. Otherwise the programme trains itself to one analyst’s unreviewed preference.

An assurance statement with narrow claims

A release record should say exactly what has been demonstrated. For example: selected decision types compile case scope, purpose, effective policy and evidence classes; deterministic fixtures showed no forbidden admission; historical policy-transition cases selected the expected version; sampled packets were reperformed by authorised reviewers; residual limitations include current-only external sources and untested document classes.

It should not say the agent is compliant, unbiased or hallucination-free. Those are not properties established by retrieval tests. It should name the period, corpus, policy and software versions covered by the evidence.

Narrow claims age better. When a new legal entity, product or source enters scope, the record makes clear that fresh fixtures and evidence contracts are needed. When an adapter changes, conformance reruns. When policy meaning changes, applicability cases are reapproved.

The context compiler is therefore not a one-time gateway. It is an operating control with source, policy, security, engineering and decision owners. Its strongest result is not a longer answer. It is a smaller, explainable evidence boundary that can be denied, replayed and challenged.

What the architecture claims: and what it does not

The design makes several testable claims. First, compiling authorisation before candidate generation can prevent a class of over-retrieval that prompt instructions cannot prevent. Second, separating applicability from relevance can reduce use of superseded or out-of-scope policy in temporal fixtures. Third, proposition-level evidence contracts can make missing support visible before answer generation. Fourth, signed plan receipts can improve reproduction of which evidence was available under which controls.

Each claim can be falsified. If a retrieval adapter ignores part of the plan, unauthorised items can still enter. If policy metadata is wrong, the compiler can confidently select the wrong rule. If the evidence contract omits a decisive proposition, the packet can look complete while being inadequate. If receipts copy sensitive content or are not retained under usable identifiers, assurance will fail.

The compiler does not guarantee truth, fairness, regulatory compliance or a sound credit decision. It does not eliminate prompt injection, insider misuse, compromised source systems or erroneous policy interpretation. It does not replace human authority. It narrows and records the evidence boundary presented to an agent.

The value is controlled selectivity: less context, admitted for stated reasons, with failure made explicit. That is a more defensible foundation for decision-grade agents than a larger vector index and a stronger instruction to “use only authorised sources.”

A compact reference architecture

The full pattern can be reduced to three planes. The governance plane owns identities, resource labels, policy scope, evidence contracts and release approvals. The compilation plane turns a request into a signed plan. The execution plane retrieves, validates and packages evidence. The language model is a consumer of the package, not a member of the trust root.

Figure 17Identity and delegation to downstream policy enforcementCausal and control schematic
Identity and delegation to downstream policy enforcement18 declared states connected by 5 authored relations. The figure supports the section A compact reference architecture. L0L1L2L3L4 01
Identity and delegation
02
Entitlement and relationship graph
03
Policy registry and applicability
04
Decision schemas and evidence contracts
05
Canonicalise request
06
Evaluate authority
07
Resolve applicable rules
08
Build and attest plan
09
Structured and graph reads
10
Filtered lexical and vector search
11
Independent candidate validation
12
Evidence manifest
13
G
14
C
15
E
16
Model drafts bounded output
17
Authorised human decision
18
Downstream policy enforcement
Boundaries: G["Governance plane · C["Compilation plane · E["Execution plane
Reading. The authored topology makes 5 declared relations across 18 states inspectable. Read it as the control structure for “A compact reference architecture”, not as measured performance. Schematic derived from the paper's authored topology; no measured quantities.

The final enforcement point still checks the human’s role and the requested action. An accurate draft does not grant authority to change a limit or continue a facility. The same principal-chain discipline used for retrieval should apply to actions.

For Northstar, this means the agent can assemble a supportable continuation assessment without becoming a credit officer, an entitlement service or a policy interpreter. Maya sees the governing evidence and the expired waiver gap. An authorised officer decides what happens next. The record preserves the plan, evidence and decision as separate artefacts.

A context window should be the end of a controlled compilation process, not the beginning of a search.