State model
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
Reviewer corrections are valuable learning signals, but an override may reflect policy exception, missing data or local preference rather than model error.
The feedback service captures the reviewed proposal, evidence, correction type and authority, then separates model defects from policy, data, process and training causes before reuse.
Organisations, systems and operating conditions are intentionally anonymised and recomposed. The design demonstrates engineering and banking-domain reasoning; it does not represent a named client estate, vendor product or measured production result.
Typed commands, state changes and receipts cross an event spine without surrendering domain ownership.
Permitted workThe assurance layer observes and evaluates; it may trigger a stop or review but cannot redefine business policy from telemetry alone.
Consistency ruleCarry one correlation chain across synchronous and asynchronous hops and separate event occurrence from observation and processing time.
Hard boundaryThe model is not a system of record, identity provider, policy authority or proof that an external effect occurred.
Agents, prompts, tools, policies and models must be built, certified, operated, changed and retired as one reachable system.
No single override updates production behaviour; feedback must be labelled, quality-checked and approved for its intended learning use.
A deterministic outer workflow contains model-led work inside typed, observable calls. Dashed messages remain proposals until policy or a human grants authority.
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
Build and runtime boundaries use the same schemas. Registry metadata drives discovery, policy, telemetry attribution and deprecation.
The deployed artefact, registry entry and evidence pack share one immutable release identifier; partial promotion is not a valid state.
These roles are deliberately vendor-neutral. Each can be independently owned, versioned and replaced.
Appends request, versions, policy result, model proposal, approval, action receipt, readback, correction and custody events under one correlation key.
Serves owned, audience-qualified and effective-dated content; exposes supersession, withdrawal and dependency metadata.
Runs component, route, trajectory, failure, harm and outcome tests against the versioned system manifest.
Presents claims beside evidence, alternatives, uncertainty, missing information, permitted actions and current custody.
Evaluates identity, purpose, capability, amount, risk tier and policy version; returns allow, deny, step-up or human-review with reasons.
Durable records carry provenance, authority, effect and custody without turning a transcript into an uncontrolled memory store.
Carry one correlation chain across synchronous and asynchronous hops and separate event occurrence from observation and processing time.
The selected design is not universally superior. It is the safer fit for this boundary and failure cost.
Use append-only events plus a rebuildable current-state projection.
Overwrite the case row with its latest status.
Cost acceptedReplay and storage are more complex, but point-in-time reconstruction and correction lineage remain possible.
Federate authoring while centralising lifecycle metadata, validation and serving rules.
Create one centrally authored knowledge corpus.
Cost acceptedFederation requires stronger contracts and owner discipline, but preserves domain accountability and release velocity.
Test prompts, models, tools, knowledge, policies, state transitions and human paths together.
Use a static answer-quality benchmark as the release gate.
Cost acceptedSystem evaluation takes longer and needs synthetic environments, but detects authority and recovery failures that answer scoring misses.
Keep people at irreversible, ambiguous and policy-exception points; sample lower-risk automated outcomes independently.
Require the same manual approval at every step.
Cost acceptedRisk-tiering reduces review load but needs calibrated thresholds, sampling and immediate withdrawal of authority when drift appears.
Compile stable decision logic and retain retrieval for explanation and residual ambiguity.
Ask a model to interpret the source document for every request.
Cost acceptedRule compilation needs controlled change, but creates repeatable decisions, regression tests and clear exceptions.
Retries are bounded by knowledge of business effect; unknown outcome remains visible, owned and independently reconciled.
Actual thresholds belong to accountable service owners. The design exposes the equations and observables that those owners must baseline.
gate_runtime = test_cases x routes x dependency_versionsrelease_capacity = available_environments / average_gate_durationoperating_cost = model + tools + platform + human_review + incidentsCollect the minimum diagnostic metadata, tokenise subjects and keep raw prompts or case evidence behind stricter access and retention.
A design is production-ready only when teams can prove what happened, recover it and change it safely.
trace and outcome completeness
control-bypass and false-negative review
cost and human-effort attribution
kill-switch and recovery exercise
Original proposal, cited evidence, reviewer change, reason code, authority, adjudication, remediation owner and validation result.
Domain and model owners adjudicate feedback and approve policy, data or model changes.
Automate evidence collection before automating approval. Expand self-service only when the paved route proves safer and faster than bespoke delivery.
Service, control, finance, model, data and operations owners interpret evidence and decide intervention.