State model
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
Production histories underrepresent rare failures and cannot be freely reused for repeated evaluation because of privacy and leakage risk.
The factory composes synthetic customer profiles, events and documents from constraint-based generators, then injects controlled contradictions and partial failures.
Organisations, systems and operating conditions are intentionally anonymised and recomposed. The design demonstrates engineering and banking-domain reasoning; it does not represent a named client estate, vendor product or measured production result.
A shared authority core coordinates independently owned capability cells, domain systems and operating evidence.
Permitted workThe platform supplies reusable controls and contracts. Domain teams retain business logic, data purpose and outcome ownership.
Consistency ruleTie deployed code, prompt, model, data, permissions, policy and evidence to one release manifest and compatibility graph.
Hard boundaryThe model is not a system of record, identity provider, policy authority or proof that an external effect occurred.
Agents, prompts, tools, policies and models must be built, certified, operated, changed and retired as one reachable system.
Synthetic provenance is machine-readable; generated data cannot be mistaken for a customer record or used as outcome evidence.
A deterministic outer workflow contains model-led work inside typed, observable calls. Dashed messages remain proposals until policy or a human grants authority.
A signed system manifest links code, prompt, model, knowledge, tools, permissions, policies, tests and owner. Any material dependency change creates a new release candidate.
Build and runtime boundaries use the same schemas. Registry metadata drives discovery, policy, telemetry attribution and deprecation.
The deployed artefact, registry entry and evidence pack share one immutable release identifier; partial promotion is not a valid state.
These roles are deliberately vendor-neutral. Each can be independently owned, versioned and replaced.
Runs a versioned state machine with explicit waits, deadlines, retries, compensations, human tasks and terminal states.
Runs component, route, trajectory, failure, harm and outcome tests against the versioned system manifest.
Classifies unknown outcomes, replays idempotent work, runs compensations and restores from the last verified checkpoint.
Builds time-qualified projections from source events and reconciliations without becoming the legal system of record.
Appends request, versions, policy result, model proposal, approval, action receipt, readback, correction and custody events under one correlation key.
Evaluates identity, purpose, capability, amount, risk tier and policy version; returns allow, deny, step-up or human-review with reasons.
Durable records carry provenance, authority, effect and custody without turning a transcript into an uncontrolled memory store.
Tie deployed code, prompt, model, data, permissions, policy and evidence to one release manifest and compatibility graph.
The selected design is not universally superior. It is the safer fit for this boundary and failure cost.
Keep consequential state transitions deterministic and use models inside bounded steps.
Let the model choose the complete path and recovery sequence.
Cost acceptedThe shell reduces flexibility, but makes deadlines, retries, permissions and recovery testable.
Test prompts, models, tools, knowledge, policies, state transitions and human paths together.
Use a static answer-quality benchmark as the release gate.
Cost acceptedSystem evaluation takes longer and needs synthetic environments, but detects authority and recovery failures that answer scoring misses.
Use local commits, idempotent steps, compensations and an unknown-outcome state.
Attempt one atomic transaction across independent banking systems.
Cost acceptedSagas expose temporary inconsistency and require recovery logic, but fit systems that cannot share one transaction boundary.
Use event-fed projections for scale and direct readback for consequential effects.
Fan out to all systems of record for every interaction.
Cost acceptedRead models introduce lag and reconciliation work, but reduce source load and make cross-system views feasible.
Use append-only events plus a rebuildable current-state projection.
Overwrite the case row with its latest status.
Cost acceptedReplay and storage are more complex, but point-in-time reconstruction and correction lineage remain possible.
Compile stable decision logic and retain retrieval for explanation and residual ambiguity.
Ask a model to interpret the source document for every request.
Cost acceptedRule compilation needs controlled change, but creates repeatable decisions, regression tests and clear exceptions.
Retries are bounded by knowledge of business effect; unknown outcome remains visible, owned and independently reconciled.
Actual thresholds belong to accountable service owners. The design exposes the equations and observables that those owners must baseline.
gate_runtime = test_cases x routes x dependency_versionsrelease_capacity = available_environments / average_gate_durationoperating_cost = model + tools + platform + human_review + incidentsEnforce purpose, tenant and workload identity at every hop; telemetry stores metadata rather than unrestricted prompts or payloads.
A design is production-ready only when teams can prove what happened, recover it and change it safely.
reachable-authority analysis
cross-version compatibility
failure and rollback rehearsal
cost, latency and trace attribution
Scenario seed, constraint set, coverage map, failure injects, expected invariants, run results and gap analysis.
Domain, risk and test owners approve coverage and acceptance criteria.
Automate evidence collection before automating approval. Expand self-service only when the paved route proves safer and faster than bespoke delivery.
Platform, domain product, data, security, model-risk and service owners share certification decisions.