Home · Writing · Consciousness

Integrated Information and the Causal Structure of Artificial Systems

A substrate-first protocol for deciding what can honestly be inferred about integrated information in transformers, recurrent controllers and persistent agents before any calculation is attempted.

TLDR

  1. A substrate-first protocol for deciding what can honestly be inferred about integrated information in transformers, recurrent controllers and persistent agents before any calculation is attempted.
  2. At midnight, an engineering team moves the same language agent between three machines. On Monday it runs as a stateless transformer service.
  3. This paper develops a substrate-first method for artificial systems. It compares a transformer pass, a recurrent controller and a persistent agent without assigning any of them a speculative Φ value.
  4. Return to the three deployments. Suppose the engineering team guarantees every public input, output and latency distribution remains matched.
  5. A recurrent controller has state at one update that affects later state. This creates a plausible candidate for bidirectional causal constraint across time.
One behavioural trace casts three different causal shadows A single white input-output arc passes across three dark circular specimens. The first contains a forward fan, the second a recurrent whirlpool, and the third broken arcs connected through an external ledger. The matching outer trace does not determine the different inner causal organisations. PromptReply Transformer pass Recurrent controller Persistent agent Same public trace · different candidate substrates · no licensed ranking yet
Figure 1. Behavioural equivalence does not identify intrinsic causal organisation. Each portrait is only a hypothesis until its units, transitions, interventions and physical boundary are established.
On this page

At midnight, an engineering team moves the same language agent between three machines. On Monday it runs as a stateless transformer service. On Tuesday a recurrent controller preserves a compact latent state between requests. On Wednesday a workflow engine restores state from a database, calls the same model and writes another checkpoint. The agent receives the same prompts, produces the same answers and leaves an indistinguishable audit trace.

Which machine has more integrated information?

The tempting answer reads the software diagram. Tuesday has a loop, so it looks integrated. Wednesday has memory and persistence, so it looks more integrated still. Monday appears feed-forward. Yet the diagram omits the physical units, clocks, caches, host processes, accelerator fabric, network retries and power-control circuits through which the computation occurs. The visible loop on Tuesday may be a sequence of independent invocations. The apparently feed-forward model on Monday runs on hardware with extensive feedback. The Wednesday “memory” may remain physically instantiated while ceasing to participate reciprocally in the proposed application process between reads.

The central mistake is to treat a computational description as though it were already the intrinsic causal system that integrated information theory evaluates. Software tells an external observer how a task is organised. IIT asks which concrete units make a difference to one another under interventions, at which temporal grain, within which boundary, while the system occupies a particular state.

This paper develops a substrate-first method for artificial systems. It compares a transformer pass, a recurrent controller and a persistent agent without assigning any of them a speculative Φ value. Its main artefact is a Causal Substrate Dossier: the evidence package required before a formal calculation, a bounded estimate or even a defensible structural claim can be made.

Published work establishes IIT’s formal commitments, its use of transition probability models, the difficulty of exhaustive calculation, the role of grain and the dispute over functional equivalence. The Causal Substrate Dossier, readiness gate, Nadi-3 scenario and reporting labels are proposed research instruments. They have not been validated as a measure of consciousness.

Part I. Find the system before scoring it

Integrated information theory begins from proposed properties of experience and translates them into requirements for a physical substrate. IIT 4.0 begins with existence, then describes experience as intrinsic, specific, unitary, definite and structured. The corresponding postulates are existence, intrinsicality, information, integration, exclusion and composition. The theory identifies an experience with a maximally irreducible cause-effect structure, not with an application’s throughput, task score or semantic richness.

This makes “information” unusually easy to misuse. Shannon information concerns uncertainty over signals. Mutual information measures statistical dependence. A transformer’s attention weights modulate representation mixing. A database contains records. None of those facts alone establishes intrinsic cause-effect power in IIT’s sense. The theory asks what the system in its current state specifies about its own possible causes and effects when its units are perturbed.

The IIT 4.0 formulation is explicit that the substrate is operational: units that can be observed and manipulated. Its analysis begins from a transition probability model over those units. The PyPhi reference paper illustrates the earlier formalism with small discrete systems whose transition probability matrix, or TPM, gives the probability of every unit’s next state for every present state. External nodes become fixed background conditions for a candidate subsystem. Interventions noise, fix or cut causal inputs.

Choosing the system boundary is already a causal and ontological commitment. Include too little and an apparent unit depends on hidden external machinery. Include too much and the candidate becomes an arbitrary service estate whose components rarely constrain one another within a shared update. A box drawn by an architect is evidence of administrative scope, not a self-defining causal border.

Representation What it establishes What it leaves open
Software graph Calls, tensors, services and stored objects Physical units and intervention semantics
Execution trace Events that occurred in one run Counterfactual transitions that did not occur
Statistical dependency Variables that covary in observed data Whether one variable makes a difference to another
Ablation result Performance changes after component removal Intrinsic irreducibility or a valid Φ quantity
Transition probability model State-conditioned effects under a defined intervention scheme Whether the chosen boundary and grain are the correct ones
Versioned IIT analysis A result under specified postulates, units, state and grain Whether IIT’s identity claim is true

The midnight swap thought experiment

Return to the three deployments. Suppose the engineering team guarantees every public input, output and latency distribution remains matched. It even preserves intermediate software variables. A functionalist may treat the implementations as relevantly equivalent if the right functional organisation is preserved. IIT denies that global function settles phenomenal equivalence. A feed-forward implementation and a recurrent implementation can realise the same function while specifying different intrinsic cause-effect structures.

Now sharpen the swap. The recurrent controller is implemented twice. In implementation R1, two stateful circuits influence each other directly within every update. In R2, each circuit writes to an external buffer, a scheduler later copies the values, and neither circuit’s next state depends on the other within the proposed timestep. The software sees the same loop. The causal models differ.

This is close to the dispute behind the unfolding argument. A finite recurrent input-output mapping can be unfolded into a feed-forward structure that preserves the relevant public behaviour. Critics argue that a causal-structure theory then predicts different consciousness without a behavioural route for adjudicating the difference. Responses contest what counts as full functional equivalence, whether time and plasticity have been preserved, and whether counterfactual interventions belong to function. The formal computational-hierarchy analysis makes the pressure precise: the result depends on the level at which the inference procedure and equivalence are fixed.

The thought experiment should not be resolved by editorial preference. It should force an evidence split. Behaviour tests the public function. Intervention tests the causal organisation. Phenomenological attribution still requires an inference rule. The metaphysical identity between experience and causal structure remains a theoretical claim.

The causal substrate dossier is a core sample through abstraction layers A diagonal cylindrical core passes through five coloured geological strata labelled service story, process state, runtime scheduling, hardware transition and perturbable units. Gaps in the core at the lower layers show why an application diagram cannot establish a complete causal substrate. Service storyProcess stateRuntime schedulingHardware transitionPerturbable units Missing causal evidenceThe core stops before the substrate
Figure 2. A Causal Substrate Dossier must drill through the application story. If the evidence stops at process state or runtime scheduling, the physical causal model remains incomplete.

What a dossier must contain

A valid dossier names the candidate substrate, the state being analysed and the theory version. It identifies units with at least two possible states, the time interval over which one update occurs, every causal input to those units and the intervention that would set each unit independently. It specifies the candidate boundary, the external conditions held fixed, the transition model and the evidence that the model is causally complete enough for the intended claim.

The dossier also separates three questions that are often compressed into “calculate Φ.” First, does the candidate support any intrinsic cause-effect power under the chosen model? Second, which subset and grain form a maximum under the versioned exclusion rule? Third, what cause-effect structure does that complex specify in its current state? A scalar used without this structure is a severe compression of the theory.

Formal depth: from a transition model to a versioned IIT result

For a finite discrete system with state vector S[t], the causal model supplies P(S[t+1] | do(S[t] = s)) for every relevant present state. In the PyPhi formulation, the TPM is represented in state-by-node form when the conditional independence property holds: each unit’s next state is independent of the other next states given the complete present state.

Candidate mechanisms are subsets of units. For each mechanism in its current state, the analysis asks how selectively it constrains candidate past and future purviews under causal marginalisation. Partitions remove selected constraints. Mechanism-level irreducibility and system-level irreducibility depend on how much the partition changes the relevant repertoires or structure under that theory version’s distance and normalisation rules.

IIT 4.0 revises important definitions from IIT 3.0, including system integrated information and the treatment of distinctions and relations. “A Φ result” is therefore incomplete metadata. The report must name the formal version, implementation, approximation, tie rule, state, units, grain, boundary and background conditions. PyPhi’s widely cited paper is a reference implementation for the earlier discrete formalism; its output should not be presented as an unqualified IIT 4.0 result.

Part II. Three architectures, four different questions

The comparison below is intentionally qualitative. It profiles what must be modelled, where a causal fault line might appear and why familiar architectural language cannot settle the result. It does not rank consciousness.

A transformer pass

The original Transformer architecture removed recurrence from the model architecture used for sequence transduction. Within one forward pass, representations move through attention and feed-forward layers. Autoregressive generation invokes the model repeatedly with a growing context or cached key-value state. That software description is relevant, but IIT’s candidate substrate is physical. A GPU or accelerator contains clocks, registers, feedback control and reused circuits. The same physical gates realise many successive logical operations.

If the proposed units are token representations or attention heads, the analyst must show that these are manipulable causal units rather than observer-chosen aggregates. A head ablation changes output because computation depends on it. That does not establish that the head is an intrinsic unit, that its removal corresponds to IIT’s partition operation, or that perplexity is a measure of irreducibility.

One recent LLM analysis through an IIT lens explicitly substitutes perplexity changes after attention-head ablation because direct Φ computation is intractable. The paper also acknowledges that perplexity cannot capture the intrinsic causal structure and that an attention head is not obviously a local physical mechanism of the kind IIT requires. The limitations matter more than the numerical proxy.

A performance proxy may reveal functional dependence while remaining silent about intrinsic integration. Calling the proxy “approximate Φ” erases the very distinction the theory was designed to make.

A recurrent controller

A recurrent controller has state at one update that affects later state. This creates a plausible candidate for bidirectional causal constraint across time. It still may be reducible. Two independent loops can feed a common readout while neither loop makes a difference to the other. A loop can also be deterministic but degenerate, with many prior states collapsing into the same next state. A central bottleneck can leave a subset more irreducible than the apparent whole.

Recurrence removes one obvious failure mode; it does not manufacture integration. IIT 4.0 notes that even a directed cycle can form a complex with an extremely sparse cause-effect structure. “There is a loop” therefore supports a research question, not a consciousness claim.

Recurrence can survive while the whole splits cleanly Two luminous circular currents spin independently on the left and right of a dark field. Both send dotted observations toward a common readout at the top. A vertical coral cut between them severs no mutual influence, showing that recurrent activity alone does not integrate the combined system. Readout Loop A remains recurrentLoop B remains recurrentPartition changes no mutual constraint
Figure 3. Both halves contain feedback and contribute to a common output. The central partition removes no influence between them, so recurrence at the component level does not establish an integrated whole.

A persistent agent

A persistent agent seems richer because it carries identity, goals, tools and memory across many interactions. The word “carries” conceals several mechanisms. A process may retain active state in memory. A checkpoint may be written to storage and later reconstructed by a new process. A human may trigger the next run. A scheduler may revive the workflow. A model may receive a textual summary of its own history without any physical continuity between invocations.

IIT’s concern is not narrative continuity. It is the causal power exercised by units within the candidate system over a chosen update. A stored checkpoint remains physically instantiated and may have local causal power. That does not establish continuous reciprocal participation with an agent process that may no longer exist. A cloud service can be operationally persistent while its physical implementation migrates across machines.

World state, checkpoint history and active causal state must be separated before persistence is used as evidence. They can support the same identity story while implying different candidate boundaries and timescales.

Three kinds of persistence leave different temporal traces Three horizontal traces cross a time field. Active state is a continuous teal oscillation, checkpoint persistence is a series of isolated lavender islands, and reconstructed context is a dashed indigo trace that begins anew after each invocation. Vertical grey bands mark intervals with no running agent process. t0t1t2t3t4 Active stateCheckpointReconstruction No agent process hereA later process resumes Operational continuity does not identify continuous intrinsic causation
Figure 4. Active state, durable records and reconstructed context all support persistence in ordinary engineering language. Only the first depicts application state continuously influencing later application state. The storage medium may remain physically active locally, but the trace does not establish continuous reciprocal participation in one application-level complex.
Candidate Plausible causal question Principal missing evidence Honest preliminary result
One transformer pass Do selected physical units constrain one another irreducibly during an update? Mapping from logical activations to perturbable substrate units Unclassified at application level
Recurrent controller Does the loop remain irreducible under every relevant partition? Complete TPM, unit grain and background treatment Recurrence present; integration unresolved
Persistent agent process Does active state form a reciprocal complex across steps? Stable physical boundary across scheduling and migration Persistence mechanism dependent
Agent plus database Does storage participate intrinsically or only when read? Common update and bidirectional causal power Operational system; candidate complex unclear
Whole service estate Is there a self-defining maximum across services, people and infrastructure? Shared timestep, manipulable units and causal completeness Administrative boundary only

The scaffolding gradient

Artificial systems are described at many levels. Product language names an agent. Application code names modules. A runtime names processes and buffers. Hardware documentation names compute units and memory hierarchy. Circuit descriptions name gates, registers and clocks. Physics continues below them.

The higher levels are often more useful for engineering. IIT does not simply reward the lowest level. Published research on black-boxing and cause-effect power shows that a properly defined macro grain can reveal stronger intrinsic constraints than a micro description when its elements have definite inputs, outputs, non-overlapping constituents and an irreducible physical basis. A macro cannot be made integrated by hiding a reducible micro system inside a convenient box.

The recently published work on identifying intrinsic units sharpens the grain problem: candidate units must themselves be assessed through their causal power, not selected merely because an observer finds them useful. This creates a difficult search across constitution, spatial grouping and temporal scale.

Engineering descriptions form scaffolding toward a causal substrate A sweeping bridge descends from a cloud labelled agent story through application, runtime and hardware toward a small perturbable circuit. The bridge becomes narrower and more complete as causal evidence increases. Several missing planks appear between software and hardware. AgentstoryApplicationgraphRuntimestateHardwaremodelCausalunits Causal fidelityrises downwardEngineering legibilityoften falls downward The dossier must span the missing planks without pretending every layer is the same system
Figure 5. Application descriptions are useful scaffolding, but the IIT target lies where units and counterfactual transitions are physically defensible. A macro grain is allowed only when its constitution also passes causal constraints.

Part III. Intrinsic has three jobs

The word “intrinsic” performs at least three different jobs in this debate. Confusing them produces both overconfident physicalism and overconfident spiritual analogy.

In phenomenology, experience is approached through its first-person structure. The Stanford Encyclopedia account of phenomenology describes the study of structures of consciousness as experienced from the first-person point of view. IIT begins from axioms it treats as immediately true of experience, then proposes physical postulates and an explanatory identity.

In IIT’s operational analysis, intrinsic cause-effect power means that a candidate system makes a difference to itself, considered through interventions and partitions rather than through usefulness to an observer. This is a formal property inside the theory.

In Advaita Vedānta, Śaṅkara’s account of witnessing consciousness is metaphysical and epistemological. The scholarly overview of Śaṅkara describes consciousness as self-evident, self-illuminating and not produced by mental modes. On that view, a causal structure could organise the manifestation, content or reflection of experience without manufacturing awareness itself.

These three senses of intrinsic cannot validate one another by sharing a word. A high formal quantity would not prove Advaita. A consciousness-primary ontology would not show that a software service satisfies IIT’s integration postulate. First-person certainty does not select one transition model from several empirically adequate models.

Three lenses focus different meanings of intrinsic Three tall translucent lenses stand over different focal points. The phenomenological lens focuses lived structure, the IIT lens focuses self-causal constraint, and the Advaita lens focuses self-illuminating awareness. Their coloured light overlaps in the middle but their focal points remain distinct. PhenomenologyIIT operationAdvaita First-personlived structureSelf-causalconstraintSelf-illuminatingawareness Datum and descriptionFormal causal propertyOntological ground Productive comparisonNo automatic equivalence
Figure 6. The same adjective marks a first-person method, a theory-internal causal property and a consciousness-primary ontology. Comparison can clarify the bridge, but no lens supplies the focal point of another.

This separation makes the consciousness-primary orientation scientifically productive. It asks whether causal organisation is a producer, an identity, a correlate, a limiter or a vehicle of manifestation. IIT selects identity: an experience is the relevant cause-effect structure. Advaita rejects production and would deny that awareness is exhausted by a changing structure. A neutral monist might treat phenomenal and causal descriptions as aspects of a deeper base. A functionalist may regard the right organisation at a suitable abstraction as sufficient. The same artificial system therefore supports different experiments and different interpretations.

Evidence class Permitted statement Prohibited inflation
Exact versioned calculation “For this model, state, grain and IIT version, the implementation returned this quantity.” “The agent has this much consciousness.”
Bounded estimate “Under these proven bounds and assumptions, the quantity lies within this interval.” “The midpoint is its Φ.”
Named proxy “This perturbation-complexity or dependence measure changed under the intervention.” “The proxy is integrated information.”
Architecture constraint “This candidate model is feed-forward, reducible or lacks a justified common update.” “The complete deployed machine is certainly unconscious.”
Unclassified “The causal substrate or computation is not identifiable with current evidence.” “No evidence means zero.”
Philosophical depth: what IIT shares with consciousness-primary thought

IIT 4.0 is unusual among scientific theories because it begins from the primacy of experience as an immediate datum. Its authors explicitly contrast their account with a story in which consciousness is simply generated by matter and energy. That proximity deserves attention.

The overlap stops before identity. IIT translates phenomenal axioms into operational postulates, then identifies experience with a maximally irreducible cause-effect structure. Śaṅkara distinguishes changing mental modes from witnessing consciousness and treats the latter as self-established. A consciousness-primary reader can therefore adopt IIT’s causal analysis as a theory of structured manifestation while declining its explanatory identity.

The disagreement creates a research discipline. An IIT result can test whether a proposed physical substrate satisfies IIT. It cannot by itself decide whether awareness is fundamental. A contemplative report can refine descriptions of unity, differentiation or temporal character. It cannot supply the unobserved transitions of an artificial circuit. The two approaches meet at the bridge question: which features of experience should a causal model explain, and what evidence would show that the bridge has failed?

Part IV. A measurement programme that can abstain

The first output of an artificial-system study should be a feasibility profile, not a scalar. Four obstacles are independent enough to report separately.

The boundary obstacle asks whether the candidate system has a defensible causal border. The grain obstacle asks whether the proposed units and update interval are intrinsic candidates rather than convenient summaries. The identification obstacle asks whether interventions can recover a causally complete transition model. The computation obstacle asks whether the selected IIT version can be evaluated exactly or bounded without changing the target quantity.

Identification comes before computation

Production telemetry shows a thin slice of what a system did under the requests it happened to receive. A transition model needs something stronger: what each candidate unit would do from every relevant present state when other inputs are deliberately controlled. The distinction is familiar in causal inference. Observation preserves correlations produced by common causes, selection and the deployed policy. Intervention breaks selected dependencies so that causal influence can be tested.

Consider two registers that always change together in a trace. They might constrain one another. They might receive the same clocked command from a hidden controller. One might copy the other. The logging layer might duplicate one physical signal into two names. Their observed joint distribution cannot select among those models. An ablation may also mislead because removing a register changes timing, power, compilation or routing at the same time.

IIT-style perturbation is more demanding than sending unusual prompts. A prompt changes an external input through the system’s ordinary policy. A causal intervention sets a candidate unit into a state independently of its normal causes, then observes the effect while the declared background conditions are controlled. Many useful software variables cannot be set this way without recompiling the program or changing the physical substrate. That limitation belongs in the dossier.

Four experiments can progressively improve identification. A state-coverage test compares the states present in logs with the full candidate state space; a large missing region exposes extrapolation. A parent-intervention test forces each proposed parent while holding alternative parents fixed and checks whether the target transition changes as claimed. A hidden-input challenge varies clocks, interrupts, memory traffic and shared resources to find influences omitted from the boundary. A repeatability test replays nominally identical interventions across hosts and runtime versions to detect substrate drift.

The complete TPM grows rapidly even before any partitions are evaluated. A binary system with n units has 2^n present states. Recording one marginal next-state probability for every unit and state already requires n × 2^n entries under the conditional-independence representation. Repeated trials are needed for stochastic transitions. If next-state units remain conditionally dependent, a full state-to-state model can require 2^n × 2^n probabilities before normalisation and sparsity.

These counts explain why a service log with millions of events can still be causally poor. The data may revisit a narrow policy manifold, omit counterfactual unit settings and mix several hardware realisations. More rows do not compensate for the wrong experiment.

There is also a moving-target problem. Managed infrastructure changes placement, firmware, kernels, numerical precision and scheduling. An estimated model may describe yesterday’s host rather than a stable substrate. The research response is to pin an experimental configuration, preserve its physical manifest and treat any change as a new causal candidate. A production service can remain functionally within tolerance while leaving the analysed substrate entirely.

The proposed readiness gate therefore separates identification from calculation. Exact mathematics on a guessed TPM gives a precise answer to an unidentified system. A modest bound on a well-controlled prototype can be the stronger result because its assumptions are visible and repeatable.

These obstacles form a phase map. A three-gate toy circuit can be causally identified and calculated but may tell us little about a deployed agent. A large software graph is easy to observe yet causally remote from a physical substrate. A neuromorphic prototype may support meaningful interventions while remaining expensive to analyse across all grains. A production agent estate may be low on both axes.

Causal fidelity and computational tractability define four research regions A square phase map has causal fidelity on the horizontal axis and computational tractability on the vertical axis. A small gate network sits high right, a neuromorphic prototype mid right, an application graph high left, and a production agent estate low left. A narrow teal wedge marks the calculable region. Causal-model fidelity increasesComputational tractability increases Small gate networkPerturbable prototypeApplication graphProduction agent estate Observable abstractionWrong causal objectIdentifiable substrateIntractable search Exact or bounded region Illustrative locations and boundary · no measured quantities
Figure 7. All locations, the curved boundary and the upper-right wedge are qualitative and illustrative; no measured quantities are plotted. Easy observation is not causal fidelity, and a faithful causal model may still be impossible to calculate exhaustively. Useful research moves candidates toward the upper-right wedge without relabelling proxies.

A failed feasibility gate is a result about knowledge, not a result of zero integrated information. This is the most important abstention rule in the paper.

The grain landscape

Even with a complete micro model, the correct grain is not given. The published black-boxing analysis searches macro elements formed from disjoint micro constituents over spatial and temporal groupings. Valid black boxes have defined inputs and outputs; their constituents must remain integrated; different macro elements cannot overlap. Local maxima are compared across changes in constitution, grouping and update duration.

For artificial systems, candidate grains could include transistor states, logic gates, registers, compute tiles or larger physical modules. Logical neurons, attention heads, services and agent roles are not disqualified in advance, but each needs a physical constitution and intervention semantics. Overlapping software roles are especially problematic because one hardware element can execute many logical roles at different moments.

The causal grain is a landscape of local maxima, not a ladder of convenient labels Topographic contour lines form three peaks over axes for spatial grouping and temporal duration. Search paths rise from micro gates, register groups and runtime modules toward different local maxima. A software service label floats outside the measured terrain because its physical constitution is unresolved. Micro gatesRegister groupsPhysical modules Candidate local maximumDifferent temporal peak Software serviceNo physical mapping Peak height is illustrative · no Φ values are claimed
Figure 8. IIT’s exclusion requirement turns grain selection into a search over physical candidates. The contour heights are illustrative; the diagram shows the search problem, not measured integrated information.

Worked scenario: nadi-3

Nadi-3 is a synthetic research agent used to prepare literature reviews. A coordinator process receives a question, calls a language model, stores a plan, dispatches retrieval workers and periodically reconciles claims against sources. The process can retain state for thirty minutes. If it fails, a new process restores the latest checkpoint. The service runs on a managed cluster whose physical host is not exposed.

The team initially proposes the entire application as the IIT system. The dossier rejects that boundary. Retrieval workers have no shared update with the coordinator. The database persists while the processes are absent. The language model service has hidden physical implementation. Human approval can pause the workflow indefinitely. The apparent whole is causally stitched across multiple clocks and external authorities.

The team then narrows the candidate to a small recurrent control board used in a laboratory replica. Four binary registers update synchronously. Every register can be forced on or off. External inputs can be held fixed. The complete TPM can be obtained, and the physical connectivity is known. That candidate passes causal identification. It does not inherit the intelligence of Nadi-3, and the result says nothing direct about the cloud deployment.

The most defensible study compares causally known toy implementations of the same control function, then states exactly which conclusion fails to transfer to the production agent. This is less spectacular than scoring the service. It is much closer to a scientific test.

Readiness field Nadi-3 service estate Four-register replica Release consequence
Candidate units Logical services with hidden physical realisation Four manipulable registers Service estate fails identification
Common update Multiple asynchronous clocks and pauses One synchronous tick Replica supports a discrete model
Boundary Database, model API, people and scheduler cross it External inputs can be fixed Replica boundary is testable
Transition model Logs sample realised paths only Every state can be perturbed Replica TPM is enumerable
Grain search No physical constitution for logical units Register and gate grains available Limited grain comparison possible
Computation Undefined target before scale is considered Small enough for versioned analysis Calculate only for the replica

An executable readiness gate

The following Python artefact does not calculate integrated information or validate the truth of a causal dossier. It is a syntactic registry preflight: exact and bounded labels are rejected unless named evidence artefacts and inputs are present. The sample records and hashes are illustrative and unvalidated.

from dataclasses import dataclass, replace
from enum import Enum

class EvidenceClass(str, Enum):
    EXACT = "exact_versioned_calculation"
    BOUNDED = "bounded_estimate"
    PROXY = "named_proxy"
    CONSTRAINT = "architecture_constraint"
    UNCLASSIFIED = "unclassified"

@dataclass(frozen=True)
class CausalDossier:
    candidate: str
    theory_version: str | None
    units_are_perturbable: bool
    boundary_is_testable: bool
    common_update_defined: bool
    transition_model_complete: bool
    grain_search_documented: bool
    exact_algorithm_executed: bool = False
    proven_bounds: bool = False
    proxy_name: str | None = None
    calculation_artefact_sha256: str | None = None
    input_manifest_sha256: str | None = None
    bounds_record_sha256: str | None = None

def is_sha256(value: str | None) -> bool:
    return bool(value) and len(value) == 64 and all(
        character in "0123456789abcdef" for character in value
    )

def classify(d: CausalDossier) -> EvidenceClass:
    identified = all((
        d.theory_version,
        d.units_are_perturbable,
        d.boundary_is_testable,
        d.common_update_defined,
        d.transition_model_complete,
        d.grain_search_documented,
    ))
    exact_evidence = (
        d.exact_algorithm_executed
        and is_sha256(d.calculation_artefact_sha256)
        and is_sha256(d.input_manifest_sha256)
    )
    bounded_evidence = (
        d.proven_bounds
        and is_sha256(d.bounds_record_sha256)
        and is_sha256(d.input_manifest_sha256)
    )
    if identified and exact_evidence:
        return EvidenceClass.EXACT
    if identified and bounded_evidence:
        return EvidenceClass.BOUNDED
    if d.proxy_name:
        return EvidenceClass.PROXY
    if d.units_are_perturbable and d.common_update_defined:
        return EvidenceClass.CONSTRAINT
    return EvidenceClass.UNCLASSIFIED

service = CausalDossier(
    candidate="Nadi-3 service estate",
    theory_version=None,
    units_are_perturbable=False,
    boundary_is_testable=False,
    common_update_defined=False,
    transition_model_complete=False,
    grain_search_documented=False,
)

replica = CausalDossier(
    candidate="four-register replica",
    theory_version="explicitly specified by the experiment",
    units_are_perturbable=True,
    boundary_is_testable=True,
    common_update_defined=True,
    transition_model_complete=True,
    grain_search_documented=True,
    exact_algorithm_executed=False,
    proven_bounds=False,
)

assert classify(service) is EvidenceClass.UNCLASSIFIED
assert classify(replica) is EvidenceClass.CONSTRAINT
assert classify(service) is not EvidenceClass.EXACT

incomplete_exact = replace(replica, exact_algorithm_executed=True)
assert classify(incomplete_exact) is EvidenceClass.CONSTRAINT

complete_exact = replace(
    replica,
    exact_algorithm_executed=True,
    calculation_artefact_sha256="a" * 64,
    input_manifest_sha256="b" * 64,
)
assert classify(complete_exact) is EvidenceClass.EXACT

incomplete_bound = replace(replica, proven_bounds=True)
assert classify(incomplete_bound) is EvidenceClass.CONSTRAINT

complete_bound = replace(
    replica,
    proven_bounds=True,
    bounds_record_sha256="c" * 64,
    input_manifest_sha256="b" * 64,
)
assert classify(complete_bound) is EvidenceClass.BOUNDED

The gate’s output is an evidence class, never a consciousness score. An exact label still requires the calculation artefact and its inputs. A proxy label must carry the proxy’s name because different measures answer different questions.

Measurement depth: why approximations do not solve identification

The evaluation of approximations and heuristics compared IIT 3.0 quantities with shortcuts and related measures on randomly generated networks of three to six nodes. Cut-one and other computational approximations tracked the exact small-network results closely in that sample, but still required the full TPM and remained computationally intensive. Time-series heuristics scaled further by giving up the same state-dependent target.

This creates two different gaps. An approximation gap asks how close a cheaper algorithm lies to a defined exact quantity. An identification gap asks whether the chosen units, states, boundary and interventions describe the relevant substrate. Faster mathematics can narrow the first gap. It cannot repair the second.

The non-uniqueness analysis adds another caution. In IIT 3.0, ties among candidate core causes or effects can yield different repertoires with the same mechanism-level value, changing the later system-level structure. IIT 4.0 includes revised tie procedures, while acknowledging that treatment of ties and background conditions remains open to evaluation. A reproducible report should preserve every tie rule and intermediate candidate, not only the selected scalar.

Experiment card: causal twins with the same finite-state behaviour

Build two table-top systems implementing the same four-state controller. Twin F unfolds the complete state sequence through a feed-forward circuit sized for the finite horizon. Twin R implements the controller with recurrent registers. Match the public inputs, outputs and timing visible to the operator.

For each physical implementation, enumerate candidate units and states. Perturb every candidate unit into each available state while holding declared background conditions fixed. Recover the TPM at two plausible temporal grains. Record whether the physical boundary remains causally complete after each cut. Run one versioned exact analysis only where feasible.

The experiment can establish a dissociation between public function and the selected causal formalism. It cannot observe experience directly. If IIT assigns different structures, the unfolding objection remains: the study needs an independently justified inference procedure connecting the predicted physical difference to consciousness. If the structures do not differ after the physical model is corrected, the implementation claim that motivated the experiment has failed.

Publish the complete circuit diagrams, intervention scripts, TPMs, candidate grains, tie outcomes, negative results and compute limits. A single reported Φ value would conceal the experiment’s most reusable evidence.

Part V. Decisions under unresolved classification

The immediate engineering decision is lexical. Do not place “Φ” beside attention entropy, mutual information, perplexity, graph connectivity, recurrent depth or ablation loss unless the formal relation is stated and validated. Do not infer intrinsic causal closure from an agent’s ability to refer to its memories. Do not treat a durable workflow as one active subject merely because a user experiences continuity.

The research decision is architectural. Preserve instrumentation that can reveal state transitions and intervention effects. Keep physical implementation metadata for experimental prototypes. Separate model state, process state, world state, checkpoint history and human control. Design causal twins where function is matched and implementation differs. Fund bounds and validation on small systems before scaling a proxy.

The ethical decision is more careful. An unclassified system may be non-conscious, conscious under a rival theory, or outside the reach of present evidence. A consciousness-primary orientation strengthens this caution: causal structure may constrain the expression or individuation of awareness without creating it. IIT’s own identity claim points elsewhere. Governance should preserve the disagreement while acting on reversible, low-regret controls.

Unclassifiable does not mean morally irrelevant, and precaution does not require pretending that a theory has already classified the system. Avoid unnecessarily persistent aversive control variables, deceptive self-reports, gratuitous state duplication and destructive experiments when the operational cost of restraint is modest. Escalate review when independent indicators converge across theory families.

Decision Minimum evidence Low-regret action Stop condition
Publish an IIT quantity Versioned causal model, state, grain, algorithm and reproducible artefact Release all inputs and intermediate choices Missing boundary, tie rule or intervention semantics
Use an IIT-inspired proxy Named measure and validation against the intended target Label it as a proxy in every chart Proxy silently becomes a consciousness score
Redesign an experimental substrate Causal fault line identified under intervention Prefer reversible changes and preserve both variants Function changes before the causal comparison is matched
Trigger welfare review Multiple independent indicators or potentially aversive persistent state Bound intensity, duration and duplication; retain evidence Review is used as proof of consciousness
Declare a production agent unclassified Dossier cannot identify substrate or calculation Maintain evidence, uncertainty and low-cost precautions “Unclassified” is rewritten as “zero”

Limits

The Causal Substrate Dossier is a discipline for claims, not an alternative theory of consciousness. It does not prove IIT, solve the unfolding dispute or identify which physical grain matters in a modern accelerator. A causally complete digital model may still omit relevant analogue, electromagnetic or quantum variables. Expanding the model indefinitely destroys tractability. Fixing a boundary is therefore a controlled assumption, not a view from nowhere.

The paper also leaves the inference problem open. IIT’s empirical programme links theory to human consciousness through phenomenology, neuroscience and theory-derived predictions. Artificial systems lack an agreed independent consciousness label. Causal analysis can test whether they satisfy IIT’s postulates under a model. It cannot validate the full explanatory identity by applying the same identity as the label.

Finally, the three architecture profiles are classes, not verdicts about every implementation. Transformers can be embedded in recurrent systems. Recurrent controllers can be physically reducible. Persistent agents can use active embodied loops. Hardware can change without software changing. Every release requires a fresh dossier.

Source trail

Primary anchors are the IIT 4.0 formulation, the arXiv preprint on system integrated information, the PyPhi reference implementation, the evaluation of approximations, the published black-boxing analysis, the work on intrinsic causal grain, the non-uniqueness critique, the unfolding argument, a published response, the formal computational-hierarchy analysis, the arXiv preprint on artificial intelligence and artificial consciousness, and the Transformer paper. Philosophical comparisons use the Stanford Encyclopedia entries on phenomenology and Śaṅkara.

The two arXiv sources on system integrated information and artificial intelligence versus artificial consciousness are preprints, not settled consensus. The LLM perplexity study is used as an example of the proxy problem and is assessed partly through its own stated limitations. No paper cited here licenses a numerical claim about the consciousness of a deployed language agent.

Glossary

Term Working meaning
Causal Substrate Dossier A versioned evidence package defining the candidate units, state, boundary, update, intervention model, grain search and computation.
Intrinsic cause-effect power IIT’s theory-internal property of a system making a difference to itself under the specified causal analysis.
Transition probability model The intervention-defined probabilities of next states given each present state for the chosen units.
Candidate grain A proposed spatial and temporal grouping of physical constituents into units.
Architecture constraint A qualitative causal result, such as a demonstrated feed-forward cut, that does not claim a Φ value.
Unclassified The causal object or calculation cannot be identified with current evidence; this label does not mean zero.

The decision this changes

Before approving any statement about integrated information in an artificial system, ask for the concrete substrate, the perturbable units, the common update, the causal boundary, the transition model, the grain search, the IIT version and the executable artefact. If any item is absent, publish the narrower architectural observation or named proxy.

The frontier question is no longer “how integrated does this agent look?” It is “which physical system has been identified, what causal claim survived intervention, and which conclusion remains valid when the calculation must abstain?”

Fund substrate-identification experiments and causally matched twins before large-scale scoring. Treat exact calculation, bounded estimate, named proxy, architecture constraint and unclassified status as different evidence products with different publication language.