State model
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
A model change can affect data, features, thresholds, users, controls and downstream decisions beyond the component being upgraded.
The assessor builds a dependency graph, compares the proposed release with the approved baseline and creates a risk-tiered evidence plan.
Organisations, systems and operating conditions are intentionally anonymised and recomposed. The design demonstrates engineering and banking-domain reasoning; it does not represent a named client estate, vendor product or measured production result.
A shared authority core coordinates independently owned capability cells, domain systems and operating evidence.
Permitted workThe system assembles evidence and tests obligations. Accountable control owners decide ratings, exceptions, findings and closure.
Consistency ruleLink each conclusion to effective obligation, control, population, test and owner; corrections append rather than erase prior evidence.
Hard boundaryThe model is not a system of record, identity provider, policy authority or proof that an external effect occurred.
Agents, prompts, tools, policies and models must be built, certified, operated, changed and retired as one reachable system.
Promotion is blocked until required tests, approvals, rollback and monitoring evidence are present.
A deterministic outer workflow contains model-led work inside typed, observable calls. Dashed messages remain proposals until policy or a human grants authority.
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
Build and runtime boundaries use the same schemas. Registry metadata drives discovery, policy, telemetry attribution and deprecation.
The deployed artefact, registry entry and evidence pack share one immutable release identifier; partial promotion is not a valid state.
These roles are deliberately vendor-neutral. Each can be independently owned, versioned and replaced.
Returns candidate entities and typed relationships with match features, contradictions, effective dates and non-merge evidence.
Evaluates identity, purpose, capability, amount, risk tier and policy version; returns allow, deny, step-up or human-review with reasons.
Appends request, versions, policy result, model proposal, approval, action receipt, readback, correction and custody events under one correlation key.
Runs component, route, trajectory, failure, harm and outcome tests against the versioned system manifest.
Durable records carry provenance, authority, effect and custody without turning a transcript into an uncontrolled memory store.
Link each conclusion to effective obligation, control, population, test and owner; corrections append rather than erase prior evidence.
The selected design is not universally superior. It is the safer fit for this boundary and failure cost.
Bias consequential journeys against false merge and retain unresolved candidates.
Automatically merge the highest-scoring candidate.
Cost acceptedMore cases require clarification, but one person's authority or risk cannot silently attach to another.
Compile stable decision logic and retain retrieval for explanation and residual ambiguity.
Ask a model to interpret the source document for every request.
Cost acceptedRule compilation needs controlled change, but creates repeatable decisions, regression tests and clear exceptions.
Use append-only events plus a rebuildable current-state projection.
Overwrite the case row with its latest status.
Cost acceptedReplay and storage are more complex, but point-in-time reconstruction and correction lineage remain possible.
Test prompts, models, tools, knowledge, policies, state transitions and human paths together.
Use a static answer-quality benchmark as the release gate.
Cost acceptedSystem evaluation takes longer and needs synthetic environments, but detects authority and recovery failures that answer scoring misses.
Retries are bounded by knowledge of business effect; unknown outcome remains visible, owned and independently reconciled.
Actual thresholds belong to accountable service owners. The design exposes the equations and observables that those owners must baseline.
gate_runtime = test_cases x routes x dependency_versionsrelease_capacity = available_environments / average_gate_durationoperating_cost = model + tools + platform + human_review + incidentsSeparate privileged, investigation and employee data from general model context; preserve legal-hold and access evidence.
A design is production-ready only when teams can prove what happened, recover it and change it safely.
obligation-to-control coverage
sampling and population integrity
evidence freshness and independence
finding closure and residual-risk approval
Dependency delta, evaluation results, population slices, control changes, approvals, release receipt and rollback proof.
Model owners and independent validators approve the release.
Automate evidence collection before automating approval. Expand self-service only when the paved route proves safer and faster than bespoke delivery.
First-line control, second-line risk, compliance, audit and legal owners keep their separate decision rights.