What would it take to run a multi-agent language model system inside a global financial institution, where every action is subject to audit and every decision must be attributable to an accountable party?
I argue that the distance between multi-agent research demonstrations and regulated production deployment is primarily a governance gap, not a capability gap. Synthesising the 2023-2026 literature on orchestration frameworks, communication topologies, failure taxonomies and evaluation, I propose a four-part reference architecture, a certainty gradient, an autonomy ledger, topology governance and adversarial twin verification, offered as synthesis and position rather than novel empirical result.
- Topology shape, not raw agent count, is the dominant predictor of task success, with returns diminishing beyond roughly a dozen agents (Qian et al., 2024)
- MAST's fourteen failure modes fall into three categories, specification, inter-agent misalignment and verification, each carrying a distinct governance implication
- Cost-controlled comparisons often narrow or eliminate the apparent advantage of complex multi-agent designs over simpler pipelines (Kapoor et al., 2024)
- Practitioner-observed authority creep: individually reasonable delegation decisions accumulate into actions no single human reviewer would have approved, motivating the autonomy ledger
- Topology governance stands in direct tension with self-optimising graph research (GPTSwarm, Automated Design of Agentic Systems), which treats autonomous rewiring of the communication graph as a desirable capability
- None of the four proposed mechanisms has been evaluated in a controlled, published sense, and a well-instrumented autonomy ledger is useless without a governance body with an explicit mandate to act on what it reveals
A four-part reference architecture for regulated multi-agent deployment: certainty gradient, autonomy ledger, topology governance and adversarial twin verification