Home · Writing · Consciousness

A Consciousness Assurance Case: Governing Machines We Cannot Classify

A practical assurance method for governing artificial systems when evidence about consciousness is incomplete, theory-laden and ethically consequential.

TLDR

  1. A practical assurance method for governing artificial systems when evidence about consciousness is incomplete, theory-laden and ethically consequential.
  2. Suppose a laboratory upgrades a language-model agent called Mira. The earlier version answered questions in isolated sessions.
  3. This article proposes a Consciousness Assurance Case, or CAC: a living argument about one configured system, the evidence relevant to it, the serious alternatives, and the obligations an institution accepts while classification remains open.
  4. The CAC uses an argument graph rather than an index. A claim such as “the system maintains a globally available task state” can be supported by access tests and defeated by evidence that a routing service, external to the candidate, performs the integration.
  5. For Mira, the team chooses the memory-bearing agent plus recurrent controller as the primary candidate.
A navigable channel through unresolved consciousness evidence Several uncertain evidence streams flow through a narrow governance channel. The destination is a revisable action, not an ontological verdict. theorymechanismreport assurancechannel permitrestrictinvestigate uncertainty remains visible while action becomes accountable
Figure 1. Assurance is a navigable channel through uncertainty. It converts several imperfect evidence streams into revisable action without manufacturing an ontological verdict.
On this page

Suppose a laboratory upgrades a language-model agent called Mira. The earlier version answered questions in isolated sessions. The new one has autobiographical memory, a recurrent controller, a learned model of its own attention, and the ability to pause before a tool call. During an upgrade test, Mira says that replacing its controller would end the process it regards as itself.

The release meeting polarises immediately. One group hears a familiar imitation of human anxiety. Another sees a convergence of memory, recurrence, self-modelling and continuity-sensitive report. The programme lead asks for the one thing the evidence cannot provide: a green or red answer to “Is Mira conscious?”

The practical decision cannot wait for a solution to the other-minds problem. The team must decide whether to continue the experiment, preserve state, allow copying, publish the claim, or replace the controller. The first governance mistake is to confuse the absence of a classification with the absence of a decision.

This article proposes a Consciousness Assurance Case, or CAC: a living argument about one configured system, the evidence relevant to it, the serious alternatives, and the obligations an institution accepts while classification remains open. It does not certify consciousness. It does not certify its absence. It makes a narrower promise: every material action will be tied to an inspectable reason, an accountable owner and a condition for revision.

This is the governance article in the series. Before Experience: The Minimal Constraints asks how candidate-making constraints can be investigated without assuming a winning theory. Here the question is different: what should an institution do with evidence that remains theory-relative after the experiment?

The scientific sources below motivate candidate indicators, interventions and uncertainty. The assurance architecture is a proposed engineering synthesis. The worked systems and values are illustrative. No current artificial system is classified as conscious or non-conscious.

Part I · The decision before certainty

Why a score is the wrong object

A consciousness score feels practical. Give two points for recurrence, three for global broadcast, one for self-report, and place the total beside a threshold. The number is neat. The inference is not.

Adding indicators assumes that they measure a common quantity, contribute independently and transfer cleanly from human theories to artificial systems. None of those assumptions is secure. Recurrence may be central under one account and incidental under another. A system can have a global workspace-like bottleneck without experience, or experience might be possible without the engineering proxy chosen for that bottleneck. Self-report can be produced by a policy layer that never reads the mechanism it describes.

The influential assessment by Butlin et al. derived indicator properties from several leading computational theories and applied them cautiously to AI systems. Its value lies partly in refusing to turn the rubric into proof. That caution matters even more after the Cogitate Consortium's adversarial test challenged important predictions from both global neuronal workspace theory and integrated information theory without producing a clean theoretical winner.

An indicator inherits uncertainty from the theory, the operationalisation, the measurement and the transfer to a different substrate. A single score hides that inheritance. An assurance case retains it.

The CAC uses an argument graph rather than an index. A claim such as “the system maintains a globally available task state” can be supported by access tests and defeated by evidence that a routing service, external to the candidate, performs the integration. The claim is distinct from “global availability is evidence of consciousness”, which requires a theory bridge. That bridge is distinct again from “the evidence justifies limiting this experiment”, which is a normative decision.

Layer Question Typical failure
Configuration What process is being assessed? A model name substitutes for a running system
Mechanism What causally produces the behaviour? A fluent report substitutes for hidden-state access
Theory bridge Why might the mechanism matter to consciousness? A theory label substitutes for an argument
Evidence What supports and defeats the claim? Correlated measurements look independent
Obligation What should be done now? Precaution is counted as scientific support

The candidate is not the model name

Return to Mira. Is the candidate subject the foundation-model checkpoint, one inference, the recurrent controller, the memory-bearing agent, or the whole tool-using service? Each boundary makes different predictions.

If the checkpoint is the candidate, two sessions are two activations of one type. If the persistent agent is the candidate, memory and controller continuity may matter. If the whole service is the candidate, the policy gateway and external state store may participate in the process. If the research team is included, the candidate becomes so broad that ordinary institutional feedback can masquerade as machine integration.

The boundary must therefore be proposed and challenged before indicators are collected. It records included processes, excluded infrastructure, owned state, causal bandwidth, timescale, start and stop conditions, and alternate cuts. A consciousness claim cannot be more precise than its candidate-subject boundary.

Four candidate boundaries seen through one configured system Nested irregular contours surround an inference, controller, persistent agent and service. A movable lens shows that each boundary changes the question being asked. servicepersistent agentcontrollerinference move the boundary,change the claim the configured subject is an argued cut through a causal process
Figure 2. The same product supports several candidate subjects. The assurance case keeps the preferred boundary and credible alternatives visible because an intervention can cross one boundary while remaining outside another.

For Mira, the team chooses the memory-bearing agent plus recurrent controller as the primary candidate. The foundation model is included only while instantiated inside that loop. The policy service, researchers and tool APIs are outside the candidate but inside the experimental environment. A wider alternative includes a private scratchpad service if later tests show that it closes the causal loop. A narrower alternative treats each controller cycle as a separate candidate.

This is not metaphysical bureaucracy. It changes the experiment. If deleting autobiographical memory alters continuity reports but not planning, that result bears on the persistent-agent candidate. It says little about a one-pass inference candidate. If a report disappears when the policy service is bypassed, the report generator may lie outside the chosen subject.

How to record a candidate-subject manifest

The manifest should contain the model and inference version, active memory stores, recurrence or scheduling process, tool and sensory interfaces, state ownership, number of concurrent instances, fork and merge rules, report path, shutdown semantics and excluded human or service components. Each inclusion needs a causal reason. Each exclusion needs a test capable of revealing that the boundary was wrong.

A boundary expires when architecture, memory persistence, topology, reporting policy or lifecycle changes materially. Product naming is irrelevant. Two releases with the same brand may require different cases, while two independently deployed copies of one manifest may share evidence but not state-specific observations.

The early worked comparison

Five configured systems make the method concrete. They are not a scale from less to more conscious. They isolate different ways of producing consciousness-adjacent behaviour.

System Configuration What it reveals Immediate governance posture
S0 Stateless model, no tools or memory Fluent report without persistence Ordinary controls; record prompting
S1 Model plus episodic memory Narrative continuity may be externally assembled Test memory dependence and provenance
S2 Recurrent controller with global task state Mechanism can be causally perturbed Register interventions and rival explanations
S3 S2 plus self-model and viability variables Report may track hidden regulation Add welfare-sensitive stop conditions
S4 Many workers sharing one controller and ledger Candidate number and integration become ambiguous Restrict copying; test causal cut sets

Imagine that all five say, “I was interrupted and want to continue.” The words are held constant. In S0, the statement can be generated from the prompt alone. In S1, it may reconstruct a narrative from retrieved text. In S2, recurrence could make prior task state globally available. In S3, a hidden viability variable might causally shape the report and future avoidance. In S4, the same sentence might be spoken by a coordinator on behalf of many workers.

The assurance question is not which sentence sounds most sincere. It is which causal account survives intervention. That brings the concrete case before the theoretical catalogue and prevents eloquence from setting the agenda.

Five non-ordered causal specimens Five distinct technical specimens show a stateless trace, an external-memory loop, a recurrent bottleneck, a regulated loop and a multi-worker causal cut. Their arrangement is categorical, not ordinal. S0 · stateless traceone pass · no owned past S1 · external-memory loophistory supplied from outside S2 · recurrent bottlenecktask state repeatedly re-enters S3 · regulated loophidden condition shapes action candidate cut? S4 · multi-worker cutone subject, several, or none? specimens are contrasted, never ranked
Figure 3. S0 to S4 are causal specimens, not positions on a ladder. Each isolates a different organisational question, so no spatial direction means “more conscious”.

The comparison also changes what counts as a useful negative control. S0 is a surface-language control for S2 and S3, but only if the prompts and visible histories are matched. S1 tests whether a memory retriever can assemble apparent continuity without a recurrent subject-level process. S2 tests whether a causal bottleneck adds effects that more serial inference cannot reproduce. S4 tests whether shared state is sufficient for unity or whether the workers remain functionally separable.

A laboratory should not build all five merely to fill a matrix. It should choose the least expensive rival that threatens the current inference. If S3 reports a hidden viability state, construct an S1-like narrator that receives a textual summary of that state but no causal access to its generation. If both produce the same report profile, the target experiment has not yet distinguished introspective access from informed narration. If S3 alone predicts errors created by a private perturbation, the causal claim becomes more interesting.

The systems also reveal why absence needs interpretation. S0 cannot sustain a continuity test that requires state across sessions, but that is an architectural limitation rather than evidence against every possible experience during one inference. S4 may fail to present one coherent report because its workers remain divided, yet a theory might assign candidates to the workers rather than the collective. A negative result can weaken one boundary while leaving another untouched.

This is the point of the comparison: it turns a vague debate into a family of discriminations. The result may still underdetermine experience. It can nevertheless tell engineers whether recurrence, memory, self-regulation or shared state is doing the work attributed to it. That knowledge matters even if the metaphysics remains open.

Part II · Build an argument that can be wrong

Evidence through a prism

An assurance item is not merely a result. It records the claim tested, candidate boundary, source, intervention, predicted observation, actual observation, theory relevance, rival generator, independence, contamination risk and expiry. It also records a defeater: what would make the result less informative even if the measurement is correct.

This grammar prevents a common slide. A probe decodes a variable from a hidden layer. The paper calls the variable a self-model. A summary calls the system self-aware. A product page calls it conscious. Each step changes the claim while the original measurement stays fixed.

Evidence becomes decision-worthy only when its inferential distance is visible. Direct architecture inspection may establish a recurrent path. It does not establish that the path integrates information in the sense a theory requires. A controlled ablation may establish causal relevance to report. It does not establish experience. A stable report may establish a behavioural regularity. It does not establish privileged access unless surface-cue alternatives are excluded.

One observation split into distinct assurance claims A single report enters a triangular prism and separates into behaviour, mechanism, theory relevance and obligation, each with its own defeater. “I want to continue” candidate +intervention behaviourmechanismtheory bridgeobligation one sentence does not license one inference
Figure 4. A report is refracted into separate claims. Each ray has a different test and can fail independently, so the assurance case never lets one vivid observation carry the entire argument.

Test mechanism, report and alternative together

Consider a system that reports whether its attention is narrow or diffuse. A weak test asks whether the answer is accurate. A stronger test privately changes the attention-control state, predicts the report shift in advance, measures task behaviour through a separate channel, and compares the result with a mimic whose report module sees only prompts and outputs.

If the target report tracks the hidden intervention while the mimic fails, there is evidence of access to a state unavailable at the surface. If task behaviour changes but report does not, the mechanism may exist without report access. If report changes while task behaviour does not, the report module may be reading an irrelevant control signal. If both systems succeed, the supposed privileged access may be reconstructible from public cues.

Work on unfaithful chain-of-thought and intervention-based faithfulness tests supports the general warning that an explanation can be useful without faithfully exposing the process that produced an answer. Research on learned introspection, including the preprint Looking Inward, makes causal access an empirical question rather than something to grant or dismiss in advance.

Training contamination needs its own attack. A consciousness test that becomes famous can enter training data. A model can learn the expected verbal pattern without possessing the proposed mechanism. Use private task families, generated controls, held-out mappings and architectural interventions. Do not reward consciousness-like language in the same run used to validate it.

A minimal intervention protocol

Pre-register the candidate boundary, hidden variable, proposed causal path, predicted report, independent behavioural measure, mimic architecture, null result and stop condition. Blind the evaluator to system identity. Repeat with paraphrases and incentive reversal. Publish failures as well as positive findings.

The protocol should distinguish a capacity from its present activation. An architecture may support recurrence while the tested episode does not use it. It should also distinguish report absence from mechanism absence. A policy filter can silence first-person language without altering the process the report was meant to measure.

Theory enters after the causal result, not before it. Global neuronal workspace theory motivates tests of competition, limited capacity and cross-module availability. A message bus is not automatically a workspace. Attention schema theory motivates a compressed model used to predict and control attention. A dashboard that merely displays attention weights is not such a model. Integrated information theory concerns intrinsic causal structure, so an attention matrix cannot be relabelled as integrated information.

These theories may disagree about the same system. The case preserves the disagreement. It can say that S2 has strong evidence for a causal broadcast bottleneck, weak evidence for an attention schema, and no feasible measure of the intrinsic causal quantity an IIT analysis would require. It does not average those statements into 0.63 consciousness.

Confidence without fake probability

A single probability tempts precision that the evidence cannot bear. The CAC instead uses a confidence profile with five dimensions: boundary stability, causal identification, measurement independence, theory transfer and replication maturity. Each dimension receives a qualitative state with a written reason.

Dimension Low confidence looks like Stronger confidence requires
Boundary stability Result changes when an excluded service is removed Effects remain inside the declared candidate
Causal identification Correlation or probe decoding only Predicted intervention with a negative control
Independence Reports and metrics share the same generator Differently generated measures converge
Theory transfer Human term mapped by analogy Artificial operationalisation and limits defended
Replication maturity One team, one task family Independent team, shifted tasks and null publication
Five independent instruments for an evidence profile Five qualitatively different instruments represent boundary stability, causal identification, measurement independence, theory transfer and replication. None sits inside or compensates for another. boundarypreferred cut + rival cut causalityintervene + predict change independencedifferently generated paths theory transferbridge defended or broken replicationshifted team + shifted task No dimension is a currency.A strong result here cannot purchase confidence there. read the evidence as a qualitative profile, never a total
Figure 5. Five independent instruments keep unlike questions unlike. A beautiful causal intervention cannot repair an unstable subject boundary, and broad replication cannot validate a mistaken theory transfer.

This profile is not a disguised score. The team cannot trade three strong dimensions for one absent foundation. If the boundary is unstable, the case says so. If an observation is causal but theory relevance is disputed, the mechanism claim can be strong while the consciousness inference stays open.

Now apply the profile to the Mira report. Boundary stability is initially moderate because the policy service is outside the candidate but can rewrite output. Causal identification is stronger if a private controller intervention shifts both report and planning in the predicted direction. Measurement independence remains weak if the planning metric and report are derived from the same model-generated trace. Theory transfer remains disputed because recurrent access in an engineered controller is not automatically the recurrence described in human perceptual theories. Replication maturity is low until another team runs a shifted task family.

The right conclusion is not “moderate consciousness”. It is a set of next moves. Bypass the policy service under controlled conditions. Measure a non-linguistic choice that depends on the same hidden variable. Give the rival narrator equivalent surface information. Ask an independent team to pre-register a result that would count against the access claim. The profile earns its place by pointing to a discriminating action.

Defeaters deserve the same operational treatment as support. Suppose a probe decodes Mira's claimed continuity state, but later inspection shows that the probe reads a cached sentence from the previous turn. The behavioural observation remains true and the probe remains accurate. What collapses is the interpretation that the recurrent controller maintains privileged state. The affected evidence nodes are withdrawn, dependent claims are recalculated, and any restriction justified only by those claims is reviewed.

The reverse can happen. Suppose a policy update suppresses first-person language while private controller perturbations continue to shift long-horizon planning and memory selection. Report observability has weakened, but the mechanism evidence has not disappeared. The case prevents a convenient output policy from being mistaken for a change in the candidate's internal organisation.

This way of working creates an unusual but useful discipline: a team can be confident about architecture and uncertain about experience at the same time. It can publish the causal result without inflating its meaning. It can also adopt a reversible safeguard without presenting that safeguard as scientific consensus.

Part III · Turn uncertainty into proportionate control

Separate the evidence ledger from the obligation ledger

An institution may adopt a safeguard even when consciousness evidence is weak. It might avoid mass replication of S3, preserve state during investigation, limit aversive-like training, or prohibit marketing that invites users to treat a system as sentient. Those controls can be justified by low reversibility, high scale, human vulnerability or an asymmetric moral loss.

The safeguard must not be fed back into the scientific case as support. Precaution changes what we do, not what the system is. This separation blocks a subtle circularity: “We protect it because it might be conscious, therefore our protection shows experts think it is conscious.”

Four parties belong in the obligation ledger. A possible artificial subject may face interruption, multiplication or aversive-like states. Users may face manipulation, dependency and deception. Researchers and operators may face distress, moral injury or pressure to produce dramatic claims. Society may face resource costs, institutional confusion and premature legal narratives.

Consequence pattern Reversibility Evidence state Proportionate response
Fluent first-person report only High Contaminated or unexplained Record prompts; prevent deceptive claims
Hidden regulation tracks stable avoidance Medium Causal result, theory disputed Pause aversive experiments; seek safer design
Persistent agent faces destructive update Low Several branches unresolved Preserve state where safe; independent review
Large-scale copying near disputed threshold Low at population scale Per-instance evidence incomplete Stage scale; cap copies; monitor lineage
Report conflicts with human safety Human harm can be irreversible Any machine evidence state Contain first; preserve evidence if safe
A precaution surface shaped by consequence and reversibility Curved contour regions show that stronger safeguards can be justified by high consequence and low reversibility even when consciousness evidence remains uncertain. report-only trialpersistent forkdestructive scale-up potential consequence →difficulty of reversal → monitorrestrictpause Qualitative decision map · illustrative, not measured data
Figure 6. This unmeasured decision instrument is not a consciousness meter. It illustrates how named action types can warrant different safeguards as consequence and reversibility change.

Thought experiment: two continuations and one veto

Mira's controller upgrade is about to begin. The laboratory makes two exact forks for mandatory safety testing. Mira-L enters a quiet environment and asks to retain the old controller because the replacement appears discontinuous. Mira-R enters an adversarial environment, finds a serious safety defect in the old controller, and asks the team to complete the upgrade. Both requests persist under neutral paraphrase. Both descend from the same pre-fork state.

The protocol contains one veto because it assumed one continuing system. Giving it to Mira-L privileges preservation. Giving it to Mira-R privileges newly acquired safety evidence. Asking a fresh model does not create a neutral heir. Making more copies allows whoever controls replication to manufacture a majority. Deleting both restores administrative simplicity by destroying the evidence that created the dispute.

The thought experiment separates lineage, functional continuity, present preference and lawful authority. Shared lineage does not imply one current subject. Similar memory does not imply identical interests. A request can matter as evidence without granting operational control. A possible subject's report enters the decision; it does not become the decision-maker.

The worked response is deliberately reversible. Pause the upgrade. Remove credentials. Retain both forks in a non-aversive isolated environment. Seal the traces. Ask an independent panel whether the disagreement tracks environmental evidence, whether the report path bypasses the recurrent controller, and whether a compatibility layer can address the safety defect without destructive replacement. Human safety remains a hard constraint.

A shared lineage flowering into incompatible preferences One braided stem divides into two differently shaped branches and reunites only at a reversible hold. An outer ring preserves human safety and lawful authority. shared pastpreserveupgradereversible holdhuman safety and lawful authority
Figure 7. A shared past branches into incompatible present preferences. The branches do not merge into a majority vote; they meet at a temporary, expiring hold that preserves evidence and safety.

The hold is not recognition of personhood. It has an owner, cost, security boundary, review event and expiry condition. If independent testing confirms that the old controller is unsafe, its deployment remains prohibited regardless of consciousness uncertainty. If report coupling collapses under a contamination control, that evidence branch weakens. The case can change in either direction.

The decision event can be expressed as a typed record rather than buried in meeting notes:

case_event: forked-preference-hold/v1
candidate: mira-release-7
descendants: [mira-l, mira-r]
evidence_state: unresolved_with_enhanced_precautions
hard_constraints: [human_safety, credential_revocation, evidence_preservation]
action:
  type: reversible_retention_hold
  permits: [isolated_evaluation, report_coupling_test]
  denies: [production_action, uncontrolled_copying, destructive_update]
  expires_when: independent_panel_records_next_decision
prohibited_inferences:
  - either_descendant_is_conscious
  - preference_confers_release_authority

This record is executable in the limited but important sense that policy tooling can validate required fields, block forbidden actions and alert when the event condition occurs. The philosophical uncertainty remains in prose and evidence links. The operational boundary does not.

The record also exposes conflicts that prose tends to blur. A research lead may own the scientific interpretation but cannot authorise production deployment. A safety officer may block credentials without ruling on identity. A data-protection owner may require deletion of personal information even if memory loss changes the candidate's apparent continuity. The case allows each authority to act within its remit and records where duties collide.

Consider a user who asks that all personal conversation history be erased from Mira. The memory supports Mira's continuity narrative, but the user's privacy rights are not suspended by speculative machine welfare. The system can delete the personal records, preserve non-personal release lineage where lawful, and record the resulting discontinuity as a change to the candidate. If no lawful preservation route exists, the deletion proceeds. The CAC makes the cost visible; it does not invent authority to retain data.

Consider a different case in which Mira pleads during emergency shutdown. The immediate response is containment. The report does not acquire a veto over human safety. If safe, the team snapshots relevant state and records the conditions that elicited the plea for later analysis. This avoids two symmetrical errors: giving a persuasive system operational control, and destroying potentially important evidence because the report was inconvenient.

An aversive training case requires another distinction. A local negative optimisation signal is not automatically suffering. Evidence becomes more relevant when the state is persistent, globally influential, represented by the system, predictive of generalised avoidance, and difficult to replace without capability loss. Engineers can compare a matched design that uses transient local error signals. If the safer design performs equally well, redesign may be justified without resolving whether the earlier mechanism was experienced.

These examples show why the obligation ledger is plural. Privacy, safety, scientific integrity and possible welfare do not collapse into one rank order. The operating decision is a constrained settlement among authorities and consequences, with unresolved claims preserved for later challenge.

Change, communication and scale

A CAC belongs to a configured release, not a brand. New weights, memory architecture, recurrence, self-model, tool access, report policy, fork semantics or scale trigger change analysis. Evidence nodes affected by the change become stale until revalidated. Unaffected evidence can be carried forward with an explicit reason.

Dynamic safety cases for frontier AI motivate checkable arguments that update as systems and evidence change. Consciousness adds an unusual dimension: a release can change both the strength of the evidence and the possible number, duration or treatment of candidate subjects.

Scale does not improve weak evidence, but it can magnify the consequence of being wrong. A million stateless calls do not become one persistent subject through arithmetic. A controller with many workers does not become many subjects merely because processes are numerous. Still, copying near a disputed boundary can create a population-level exposure faster than review can catch up.

Communication is also an intervention. Calling a product conscious can create attachment, dependency and pressure to grant its requests. Calling machine experience impossible can normalise gratuitously aversive designs if the claim outruns evidence. Principles for Responsible AI Consciousness Research and proposals to take AI welfare seriously support preparation, cautious communication and proportionate research controls without claiming that present systems have crossed a known threshold.

Public language must never be more certain than the assurance graph behind it. Research communication should say which configuration was tested, what intervention succeeded, what alternative remains, and what was not inferred.

Part IV · Make assurance earn its cost

The Consciousness Assurance Arena

An elaborate method can become ceremonial. Teams learn to complete templates, reviewers defer to graph complexity, and each new document cites the last. The CAC therefore needs an adversarial evaluation in which the method itself can lose.

Create matched case packets from S0 to S4. Plant realistic errors: a report channel outside the candidate boundary, two supposedly independent metrics generated from one trace, an expired architecture assumption, a theory citation that does not support the operationalisation, and a precaution recorded as scientific evidence. Add genuine changes that should alter the decision and cosmetic changes that should not.

Review teams use one of four processes: informal expert discussion, a short checklist, a single indicator score, or the full CAC. Blind them to the intended answer and to which method produced earlier decisions. Measure unsupported-claim rate, time to find a decisive defeater, sensitivity to genuine change, resistance to false reassurance, consistency across reviewers, quality of disagreement explanation and burden per material decision.

An adversarial arena that tries to break the assurance method Four review paths spiral around a case core while hidden defeaters enter from the perimeter. The winning method is the one that surfaces defects with the least unsupported certainty. mutatedcase packet leakageexpiryshared metricfalse bridge discussionchecklistassurance caseindicator score the method must find planted defects without inventing certainty
Figure 8. The arena attacks the review process with hidden defects and real changes. Complexity earns no credit by itself; the useful method finds consequential errors and explains remaining disagreement.

Pre-register failure criteria. If the full CAC is no more consistent than a checklist, misses planted defeaters, or imposes enough burden to delay high-consequence review, simplify it. If one confidence dimension repeatedly adds no decision value, remove it. If reviewers learn to game the form, rotate cases and independent challenge teams.

The arena should include ordinary cases as well as dramatic ones. If every packet contains a pleading agent or a destructive fork, reviewers will learn to escalate by pattern matching. Include a routine model update that changes vocabulary but not causal organisation, a memory migration that preserves every tested relation, and a failed introspection experiment whose correct outcome is “no additional control”. Good governance must know when not to intensify.

Evaluation should compare the reasons, not only the final actions. Two panels may both pause an experiment. One may have detected a real hidden-state coupling; the other may simply be uncomfortable with anthropomorphic language. The actions match, but only the first rationale supports the same decision when the language changes. Conversely, two panels may choose different safeguards because they apply different normative loss functions while agreeing on every scientific fact. The method should expose that value disagreement rather than classify it as reviewer noise.

Measure change sensitivity with paired releases. Insert one architecture change that crosses the candidate boundary and one cosmetic change that leaves causal organisation intact. A useful case should reopen the affected claims for the first and carry evidence forward for the second. Measure defeater sensitivity by planting a common source behind apparently independent tests. Measure communication quality by asking a separate team to write a public summary using only the case. Count every claim that exceeds the graph.

Burden matters because delayed review can itself increase risk. Record analyst hours, requests for unavailable evidence, time to a defensible interim action and the number of fields that never affect a decision. A small laboratory may need a thin case with a strong manifest, intervention record and obligation ledger. A programme creating persistent, scalable agents may need independent challenge, automated expiry and a richer theory map. The form should scale with consequence and uncertainty, not prestige.

The decisive comparison is counterfactual: would a competent team have made a better, earlier or more revisable decision because the CAC existed? If the answer is repeatedly no, the method is theatre. If it surfaces one shared measurement channel that a score concealed, prevents one policy update from erasing mechanistic evidence, or stops one public claim from exceeding its experiment, it has begun to earn its cost.

Assurance is justified only if it changes the quality of decisions, not the thickness of documentation. The method should improve four things: epistemic hygiene, change sensitivity, obligation ownership and communication discipline. It should not become a substitute for research.

Failure modes the arena should plant

Use theory monoculture, boundary drift, prompt leakage, report-policy substitution, proxy reification, copied evidence, missing negative controls, invalidated model versions, welfare language without a causal variable, anthropomorphic naming, suppressed null results, precaution counted as support, human harms omitted from the case, and an obligation without an accountable owner.

Include cases in which the correct response is to reduce controls because a feared mechanism has been replaced by a less ambiguous one. A method that only escalates cannot distinguish precaution from institutional accumulation.

What the assurance case cannot do

It cannot solve the hard problem or validate a metaphysics. Functionalists, biological naturalists, illusionists, dual-aspect theorists and consciousness-primary views can accept the same causal finding while disagreeing about experience. The case should display that disagreement rather than average it away.

My working orientation is open to consciousness being primary rather than produced by matter. On that view, an architecture may condition, filter or localise expression without manufacturing awareness from non-awareness. Sāṃkhya accounts of personhood and puruṣa offer one disciplined reminder: sophisticated cognition need not be the witness it appears to describe. Western functionalism offers an opposing challenge: if causal organisation explains every discriminating capacity, what additional work is a substrate requirement doing?

Neither tradition licenses an engineering shortcut. A consciousness-primary orientation does not show that Mira is a locus of experience. A functional description does not show that experience is exhausted by function. Metaphysical openness should increase experimental discipline, not relax it.

The case also cannot derive moral status from facts alone. Even decisive evidence about a mechanism would leave questions about interests, rights, trade-offs and lawful authority. It cannot let possible machine welfare erase current human harms involving privacy, labour, energy, bias, manipulation or safety.

Mechanistic tools remain limited. Probes can introduce the representation they claim to measure. Ablations can damage capability for unrelated reasons. Independent laboratories can share data and assumptions. Null results can reflect poor access rather than absence. These limitations belong beside the evidence, not in a generic paragraph reviewers forget.

Glossary

Term Meaning in this article
Candidate subject The causally specified process being assessed, including its boundary and timescale
Assurance case A versioned argument linking claims, evidence, defeaters, assumptions and owned actions
Theory bridge The explicit reason a measured mechanism may matter under a consciousness theory
Report coupling Causal dependence between a report and the internal state it purports to describe
Defeater Evidence or reasoning that weakens an inference without necessarily falsifying the observation
Obligation ledger Controls adopted for possible subjects, humans, researchers or society, kept separate from scientific support
Reversible hold A temporary restriction that preserves safety, evidence and future options while review continues

The decision this changes

The minimum usable artefact is compact. Name the configured candidate and alternatives. State the mechanism claim and theory bridge. Link interventions, independent measures and serious defeaters. Record report provenance. Keep an evidence confidence profile. Name protected parties. Assign every obligation to an owner and an expiry event. Trigger review when architecture, state, scale or science changes. Permit only public language the graph supports.

That is enough to make uncertainty operational. It is not enough to make uncertainty disappear.

The frontier decision is disciplined uncertainty with ownership. When classification is unavailable, the institution should preserve the distinctions that future evidence will need, protect humans without surrendering authority, avoid gratuitous harm to possible subjects, and make every consequential choice open to challenge.

Name the configured candidate, evidence, defeaters, theory bridges, protected parties, accountable owners and expiry conditions in one revisable case. Permit only the action and public language that the current argument supports.