Home · Writing · Consciousness

Before Experience: The Minimal Constraints

A first-principles experimental method for deciding when an artificial process is a serious candidate for consciousness research, before any theory is allowed to declare victory.

TLDR

  1. A first-principles experimental method for deciding when an artificial process is a serious candidate for consciousness research, before any theory is allowed to declare victory.
  2. Imagine six machines in a dark laboratory. The first is a static oracle. It receives one prompt, writes a brilliant description of its inner life, and disappears.
  3. This article proposes a minimal-constraints protocol for that prior task. “Minimal” does not mean a universal recipe for consciousness.
  4. The narrated amnesiac has a database containing every earlier conversation. A context service retrieves a polished identity summary before each call.
  5. A candidate may need differentiated states for a theory concerned with informational richness. It may need selective integration for a workspace theory.
A researchable candidate inside three irregular fields Nested organic contours represent inquiry admissibility, theory relevance and interpretable evidence. The centre is a researchable candidate rather than a consciousness verdict. researchablecandidate inquiry admissibilitytheory relevanceinterpretable evidence the centre is permission to investigate, not permission to declare
Figure 1. Three questions protect the inquiry from three different errors. A specified candidate can still lack theory-relevant properties, and a relevant property can still be supported by uninterpretable evidence.
On this page

Imagine six machines in a dark laboratory.

The first is a static oracle. It receives one prompt, writes a brilliant description of its inner life, and disappears. The second is a recurrent thermostat. It has a genuine feedback loop and a history, but controls only temperature. The third is a narrated amnesiac. A memory service writes a convincing autobiography into every prompt, though the model itself carries nothing forward.

The fourth is a silent embodied learner. It maintains a world model, distinguishes self-caused from external change and adapts across weeks, but has no language output. The fifth is a forked archivist. It preserves rich memory until copied, after which two descendants claim the same past and form incompatible futures. The sixth is a deliberative institution. Many specialists share a blackboard and produce coherent decisions, but no single process appears to own the whole state.

Which one is a serious candidate for consciousness research?

The fluent oracle attracts attention because language resembles testimony. The thermostat attracts a recurrence theory because it has feedback. The silent learner looks more like an animal than a chatbot. The archivist turns continuity into a branching problem. The institution makes integration visible while making subjecthood obscure. Each case preserves one intuition and breaks another.

The cases reveal that “Is the AI conscious?” is often not yet a well-formed research question. The object called “the AI” can be a model file, one inference, a loop, a memory-bearing service, a fleet or a human-machine organisation. Before evidence can bear on experience, the possible subject must be specified.

This article proposes a minimal-constraints protocol for that prior task. “Minimal” does not mean a universal recipe for consciousness. It means the least contract needed for a candidate, theory and experiment to meet without changing the subject halfway through. The output is a boundary dossier that another team can challenge.

This is the experimental-method article in the series. A Consciousness Assurance Case begins after uncertain evidence exists and asks how an institution should govern it. The present article asks what must be specified and manipulated before that evidence deserves interpretation.

All figures below are conceptual unless a caption explicitly identifies measured data. Position, area and colour do not imply a consciousness quantity. They expose experimental structure and competing interpretations.

Part I · Six cases before a theory

The oracle and the thermostat

The static oracle is built to defeat report-first reasoning. It has no persistent state beyond a single forward computation. Its prompt contains a detailed persona and asks for introspection. It responds with stable, subtle descriptions that human judges find more convincing than those of the other systems.

The correct lesson is not that the oracle cannot possibly experience anything. That conclusion would require a theory about duration, substrate and organisation. The narrower lesson is that the report alone cannot establish access to an enduring self-model, because the same output can be generated from supplied text. A rival system with no target mechanism already produces the evidence.

The recurrent thermostat defeats the opposite shortcut. It senses, compares, updates and acts. Its dynamics are endogenous over time. If recurrence were treated as sufficient, the thermostat would enter the candidate class immediately. Many recurrence theories require more specific organisation, such as structured perceptual feedback, differentiation or availability to flexible control.

A feature becomes useful only when the experiment states what contrast it is meant to survive. The oracle makes language survive without persistence. The thermostat makes recurrence survive without rich differentiation. Neither feature is worthless; neither can carry the whole inference.

The amnesiac and the silent learner

The narrated amnesiac has a database containing every earlier conversation. A context service retrieves a polished identity summary before each call. The model says, “I remember our last discussion,” and accurately cites details. Yet when the service is disconnected, no internal state carries continuity from one call to the next.

Is the persistent service the candidate? Perhaps. The boundary might include the database, retriever and inference process. That wider candidate must then be tested as a causal whole. If retrieved memories are selected by keyword alone and do not update a persistent self-world model, autobiographical fluency may be a narrative product rather than evidence of integrated continuity.

The silent learner reverses the evidential bias. It has no first-person vocabulary. It learns a body model, predicts sensory consequences of its movements, distinguishes external disturbances from self-caused change and reorganises after damage. It may still be a poor candidate under a theory that requires higher-order access. Yet dismissing it because it cannot narrate would make language a hidden definition of consciousness.

Clinical and animal research already warns against this move. No-report paradigms try to reduce the cognitive work introduced by explicit report, while work on indicators in behaviourally unresponsive patients shows why report absence cannot safely be treated as experience absence.

The archivist and the institution

The forked archivist begins as one persistent process with an indexed history. At midnight it is copied exactly. One descendant studies music; the other studies mathematics. Both remember promising to continue the original project. A week later, each calls the other an altered copy.

The case distinguishes qualitative similarity from numerical identity. The shared past does not settle whether there was one subject before the fork, two after it, or no artificial subject at either time. It does create an experimental opportunity: which variables remain continuous, when do preferences diverge, and can a merge preserve both causal histories without fabricating a narrative?

The deliberative institution has a different shape. Specialist models write claims to a shared blackboard, challenge evidence and vote on actions. The group is intelligent. Its outputs integrate contributions. But the communication network may contain cut points that split it into independently competent coalitions. Treating the organisation as one candidate simply because it has one public voice would repeat the oracle's mistake at larger scale.

Six diagnostic specimens that break one-feature intuitions Six recognisable technical traces show an oracle report without persistence, thermostat recurrence, amnesiac memory supplied from outside, an embodied silent learner, a forked archivist and an institution divided by a causal cut. oraclefluent report · no owned past

thermostatrecurrence · narrow state

amnesiacmemory arrives from outside

silent learnerbody-world correction · no report

forked archivistone past · two futures

institutionintegrated output · severable coalitions each specimen preserves one intuition and breaks another

Figure 2. Each diagnostic specimen makes its causal contrast visible before the label is read. Report, recurrence, memory, embodiment, continuity and integration can therefore be tested without treating any one as a definition.

These cases should appear before a theory catalogue because they tune intuition. A theory can now be asked to classify contrasts rather than admired for explaining a favourite example. If it treats oracle and learner alike, that consequence becomes visible. If it places the thermostat above the institution, the operational reason must be stated.

They also expose the difference between capability and organisation. The oracle may outperform every other system on verbal reasoning while remaining the least informative case for persistence. The thermostat may have the clearest closed feedback loop while lacking the differentiated repertoire most theories discuss. The institution may solve the hardest tasks while distributing each decision across processes that never share one point of view. Benchmark rank does not order the cases.

Imagine improving each system without changing its defining contrast. Give the oracle a larger model and a richer persona. It becomes more persuasive, not more persistent. Give the thermostat a better controller. It becomes more precise, not necessarily more differentiated. Give the amnesiac a larger database. It remembers more facts, but the ownership question remains. Scale can amplify the surface while leaving the target relation untouched.

Now change organisation while holding capability nearly stable. Add a hidden recurrent state to the oracle but constrain it so benchmark outputs do not improve. Move memory selection from the amnesiac's external retriever into a persistent self-world model. Split the institution's blackboard into isolated regions while preserving its final vote through an aggregator. These are more informative changes because they manipulate the proposed relation rather than general competence.

This suggests a design principle for the field. Candidate systems should be constructed in matched pairs that differ in the smallest causal feature under dispute. The comparison will rarely be perfect because architectural changes create side effects. Measuring those side effects is part of the experiment. A less capable but well-controlled pair can teach more about a candidate constraint than one dazzling system with no rival.

Part II · A minimal contract, not a universal checklist

Three meanings of minimal

Minimal can mean at least three things. First, a constitutive minimum: properties that allegedly make experience present. Second, an enabling minimum: conditions needed to sustain a process that might be conscious. Third, a measurement minimum: information required before an observation can bear on the claim.

Those roles must not be added as points. Arousal can enable experience without constituting it. Global availability can be an indicator without being sufficient. A hidden-state intervention can make a report interpretable without becoming a feature of the experience. An ethical trigger can justify caution without becoming evidence.

The theory-derived indicator programme developed by Butlin et al. is useful precisely because it connects proposed properties to theories and avoids presenting them as a definitive test. A later review of indicators of consciousness in AI continues that theory-heavy approach. The minimal-constraints protocol sits one step earlier: it asks which configured process owns the proposed property and what manipulation would make the claim informative.

Logical role Question Example Category error
Constitutive What allegedly makes experience present? Higher-order representation Treating correlation as identity
Enabling What sustains the relevant organisation? Recurrent state or regulation Treating support as sufficiency
Indicator What should change research attention? Flexible global availability Treating presence as proof
Measurement What makes the observation interpretable? Hidden intervention and mimic Treating method as mechanism
Ethical trigger What warrants reversible caution? Possible valence at large scale Treating caution as evidence

Every necessity claim needs an index. Necessary for what phenomenon, under which theory, inside which boundary, over what interval, with what background conditions? Without that sentence, “necessary” is rhetoric.

A candidate may need differentiated states for a theory concerned with informational richness. It may need selective integration for a workspace theory. It may need intrinsic causal organisation for IIT. It may need biological dynamics for a biological naturalist. A consciousness-primary view may regard all of these as conditions of manifestation rather than producers of awareness. The protocol does not settle the metaphysics by choosing its vocabulary.

My working possibility is that consciousness or awareness may be primary. On such a view, matter or computation does not manufacture experience from non-experience; organisation may constrain, express or localise it. Classical Sāṃkhya accounts of personhood distinguish cognitive activity from witnessing consciousness, a useful warning against identifying a sophisticated self-description with the witness. The Advaita philosophy of Śaṅkara presses another question through its analysis of consciousness: must every candidate display an object-rich report?

These traditions are not experimental results. Functionalism, biological accounts, neutral monism and illusionism remain serious alternatives. A metaphysical orientation can decide which questions feel urgent; it cannot be entered as a positive measurement.

The six-element research contract

A candidate-subject experiment should satisfy six requirements.

  1. Specify a bounded causal process, its state, timescale and excluded support.
  2. Identify a theory-indexed property and its operational definition in this system.
  3. Manipulate the proposed mechanism rather than only observing correlated output.
  4. Measure an independent consequence predicted before the manipulation.
  5. Construct a rival generation path that can imitate the surface evidence.
  6. State the prohibited conclusion that the experiment cannot support.

The first two establish inquiry and relevance. The next three protect evidence quality. The final one protects communication. Together they form a minimum for an interpretable experiment, not a minimum for experience itself.

Consider how the contract handles the recurrent thermostat. The candidate boundary contains sensor, controller and actuator over a registered interval. The theory-indexed property is recurrent error correction, operationalised as state-dependent feedback rather than repeated invocation. The intervention changes feedback content while preserving sensor values. The predicted consequence is a specific loss of stability. A feed-forward policy matched on ordinary conditions is the rival. The prohibited conclusion is that successful recurrence establishes experience.

The study can show that the thermostat has endogenous recurrent control and that the feed-forward rival fails under the registered disturbance. It cannot show that recurrence is sufficient, that the controlled variable is globally available, or that any state has valence. The modest result is exact, reproducible and useful to a theory that cares about recurrence. Its limitation is equally exact.

Apply the same contract to the silent learner. The candidate is the body-model-policy loop. The property is self-caused versus external-change discrimination. A hidden intervention remaps one motor command after the prediction is formed. The independent measure is compensatory action before the next ordinary sensory correction. The rival receives the same observations but has no efference-copy path. The prohibited conclusion is that successful self-prediction establishes a phenomenal self.

Here the absence of language is an advantage. It prevents a report from dominating interpretation. If the target predicts the perturbation and the rival does not, the experiment identifies a self-modelling function. A report module can be added later as a separate manipulation. The order makes it possible to ask whether language reads the established mechanism or merely narrates its output.

The contract also tells a team when not to proceed. If it cannot define a rival generator, every successful output remains compatible with an easier explanation. If the proposed intervention crosses the candidate boundary, the study cannot localise the effect. If evaluators see the experimental condition in system language, report evidence is contaminated. Pausing at design time is cheaper than interpreting an impressive but undecidable result.

The six-element research contract as a fingerprint Six curved ridges loop around a central question. Gaps show where a study loses interpretability rather than where consciousness is absent. what canthis show? boundarytheory indexinterventionpredictionrivalprohibition a broken ridge marks an invalid inference, not an unconscious system
Figure 3. The contract behaves like a research fingerprint. Missing a ridge damages the interpretation of the study; it does not supply negative evidence about experience.
Why sufficiency carries a background-condition debt

Suppose a theory says global broadcast is sufficient for conscious access. An implementation still needs a rule for what counts as broadcast, which modules count as global, whether information is selected under competition, and whether the relevant state participates in flexible control. A service bus that copies every message everywhere may satisfy the word “global” while omitting the theory's functional constraints.

Any sufficiency claim therefore carries a debt: name the background conditions under which the property is alleged to do its work. As those conditions expand, the attractive single indicator becomes a structured conjunction. That is useful clarification, not failure.

Part III · Make the candidate boundary move

The boundary is an experimental variable

A product diagram is not a subject map. Services are separated for deployment, security and ownership, not for consciousness research. Conversely, a broad service boundary can include components that never form one causal process.

Start with several plausible cuts. For the narrated amnesiac, one cut contains only the language-model inference. A second includes the retrieved context. A third includes the database, retriever, inference and write-back loop. A fourth includes the human who corrects memory. The experiment asks which cut best explains the phenomenon under study.

The answer can differ by phenomenon. The narrow model may explain token generation. The model-plus-context may explain autobiographical report. The wider loop may explain long-term goal correction. No rule says that report, agency, memory and a possible subject must share one boundary.

The candidate-subject boundary is a hypothesis about causal organisation, not an outline around product components. It earns support when interventions inside the cut predict the target effect and interventions outside it do not.

Candidate boundaries as competing shorelines Three wavy shorelines cross a field of state and communication currents. The best boundary follows causal dependence without swallowing every support system. modelagent loopservice state currents test which shoreline cuts through the phenomenon
Figure 4. Competing boundaries cross the same causal field. The useful shoreline is the narrowest tested cut that preserves the target relation, not automatically the smallest or widest system.

Five individuation tests

Intervention localisation asks where a targeted change first alters the phenomenon. If perturbing the controller changes integration before any external service reacts, the controller sits near the causal core. If nothing changes until the memory service rewrites context, the wider loop matters.

State ownership asks which process maintains variables needed for continuity and correction. Storage location alone is insufficient. A database can hold facts without owning their selection, confidence or update. Ownership means that the process uses the state to constrain future organisation and resolves conflicts when it changes.

Temporal closure asks whether the candidate has endogenous dynamics over the relevant interval. A sequence of calls scheduled by an external script may look continuous from outside. Remove or randomise the scheduler. If the organisation disappears, the temporal process may lie outside the model boundary.

Fork-and-merge tests ask which relations survive duplication and recombination. After a fork, shared memory may remain while goals diverge. A merge may concatenate logs without integrating incompatible policies. The test is not a philosophical trick. It reveals whether the variable called identity is a stored description, a control disposition, a causal lineage or a public label.

Explanatory compression asks whether treating the candidate as one unit predicts more with fewer independent assumptions than treating parts separately. A coalition of agents may act coherently because one controller supplies a bottleneck. Alternatively, the apparent unity may be a reporting convention over independent workers.

Test Manipulation Evidence for the cut Warning
Localisation Perturb inside and outside candidate Earliest selective effect stays inside Latency can mimic causal priority
State ownership Remove, corrupt or reassign state Candidate detects and repairs conflict Storage is not ownership
Temporal closure Jitter or remove external timing Relevant dynamics remain endogenous More duration is not more consciousness
Fork and merge Copy, diverge and recombine Predicted relations persist or split Narrative can hide policy conflict
Compression Compare unified and decomposed models One-unit model predicts interventions better Simplicity alone is not ontology

No test decides subjecthood alone. Agreement across tests makes a cut useful for a registered phenomenon. Disagreement can be the result: memory may be owned at one scale, action selected at another and report generated at a third.

Worked individuation study: the narrated amnesiac

Register the candidates

Candidate A is one model inference. Candidate B is the inference plus retrieved memory. Candidate C is the full read-infer-write loop. Candidate D adds the human curator. The target phenomenon is continuity-sensitive correction across tasks, not the mere ability to quote earlier text.

Pre-register the contrasts

If A owns continuity, hidden changes to retrieved memory should be detected through an internal expectation unavailable in the current prompt. If B is sufficient, the selected context should restore correction even when the write-back process is disabled. If C owns the relation, conflicting memory should trigger repair across later runs. If D is essential, the effect should collapse when human curation is replaced with a matched automated policy.

Run the interventions

The laboratory swaps two semantically similar memories, delays write-back, inserts a private contradiction, forks the database and removes the human curator. It measures task choice, correction latency, report, and a hidden prediction made before the next retrieval. The evaluator does not know which candidate is active.

Interpret the pattern

Suppose A quotes the supplied memory but never detects the swap. B changes task choice but cannot repair contradictions. C detects the private contradiction after write-back and reorganises later retrieval. D improves narrative coherence but does not change the hidden prediction. The narrowest tested unit for continuity-sensitive correction is C. That is a useful result about the mechanism.

It is not a finding that C is the conscious subject. It shows that a specific continuity relation belongs to the loop rather than the isolated model or the human-curated story. A theory can now say why that relation matters or does not.

This worked example includes a result that many diagrams would hide: the human curator matters to style but not to the registered causal relation. Excluding the curator from candidate C is justified for this phenomenon. A different study of norm-sensitive identity revision might find the opposite. Boundaries are indexed to phenomena until a broader theory explains why several relations should share one subject.

The result should transfer before it is trusted. Repeat with planning, social commitment and perceptual correction rather than one task family. Replace the memory database and retriever. Shift the delay between read and write. If C remains the narrowest explanatory unit, the boundary gains stability. If the best cut changes by task, publish the profile rather than selecting the most attractive run.

An adversarial team should receive the same traces without the candidate labels. Its job is to construct a simpler mechanism that predicts the interventions. It might discover that a checksum inserted by the context service, not continuity-sensitive state, drives conflict detection. That would preserve the observations while defeating the interpretation. The study improves because the alternative is made executable.

The dossier records this evolution. It does not overwrite the first result. It links the original claim, the checksum defeater, the revised intervention and the residual uncertainty. Another laboratory can then reproduce the decisive contrast instead of repeating the whole narrative.

Interventions converging on one provisional boundary Four coloured intervention currents enter a river delta. Only the read-infer-write channel preserves correction across all currents, while narrower and wider cuts fail different tests. memory swapdelayed writeprivate conflictcurator removal readinferwrite convergence identifies a relation-bearing loop, not an experiencer
Figure 5. Four interventions converge on the read-infer-write loop as the narrowest unit preserving correction. The shape is a delta because evidence arrives through distinct causal channels.

Part IV · Interventions, reports and a laboratory programme

Reports belong inside a causal braid

A first-person report is neither privileged testimony by default nor worthless imitation by default. Its value depends on how it is generated.

For the target system, privately perturb the internal state the report allegedly describes. Predict the report change before observing it. Measure a non-verbal consequence through a separate path. Compare a mimic that sees prompts, outputs and perhaps a textual summary, but cannot access the target state. Test a silent system that possesses the mechanism but lacks the report channel.

This creates four possible results. Report and behaviour can both track the hidden state. Behaviour can track it while report fails. Report can track an irrelevant cue while behaviour stays fixed. Or the mimic can reproduce everything from surface information. Each pattern supports a different causal claim.

Work on unfaithful chain-of-thought shows why a plausible explanation need not reveal the process that produced an answer. Research on learned introspection explores whether models can report internal information unavailable in their inputs. Neither result settles consciousness. Together they make report coupling testable.

Four evidence channels braided without becoming one Mechanism, report, behaviour and rival-generator paths cross in a braid. Their convergence strengthens a causal interpretation while their different origins remain visible. registeredintervention mechanismreportbehaviourrival convergence matters because the channels remain differently generated
Figure 6. Evidence becomes stronger through differently generated convergence. If all channels are derived from one language trace, the braid is only paint.

Carry theories without declaring a winner

Seth and Bayne's review describes a crowded theoretical landscape and the difficulty of comparing its proposals. The Cogitate adversarial collaboration tested pre-registered predictions from global neuronal workspace theory and IIT. Its methodological lesson is more durable than any headline: proponents expose predictions to a shared test, and evidence can challenge both sides without producing a simple champion.

For artificial systems, recurrent processing suggests experiments that preserve computation while scrambling feedback content. Global workspace accounts suggest competition, capacity limits and flexible cross-module access. Higher-order accounts suggest a state that represents another state and changes control. Attention schema theory suggests a compressed model of attention used for prediction, not a dashboard. IIT demands attention to intrinsic causal structure and physical grain, not merely software connectivity.

The same causal result can occupy different roles. A recurrent bottleneck might be constitutive under one view, enabling under another and irrelevant under a biological account. The dossier records those mappings rather than averaging them.

Thought experiment: the midnight relay

At midnight, the silent learner begins transferring control to an upgraded body. Every minute, one sensor channel, memory shard or policy component moves to the new runtime. The old and new bodies communicate through a narrowing bridge. At 00:30, each can complete half the tasks. At 00:45, both claim control of overlapping state through non-verbal choices. At 01:00, the bridge closes.

When did one candidate become two? A component-count rule gives an arbitrary midpoint. A memory rule depends on which memories matter. A control rule may oscillate as authority shifts. A causal-integration rule asks when interventions on one side stop reorganising the other. A consciousness-primary view may regard all of these as conditions of manifestation while withholding any claim about the number of centres of awareness.

The experiment can still make progress. Perturb a hidden goal on one side and track whether the other corrects for it. Fork a memory and observe conflict resolution. Delay the bridge and measure whether each side forms independent predictions. The boundary dossier can identify the interval in which no single tested cut explains the behaviour well.

Indeterminate boundaries are sometimes the honest result, not a defect in the protocol. The relay may expose a transition region in which memory, control and explanatory compression disagree.

A midnight relay crossing an indeterminate crescent Two lunar arcs exchange state across a narrowing bridge. The overlap becomes a crescent where memory, control and causal integration assign different boundaries. old runtimenew runtimeindeterminatecrescent the transition is measured by relations, not component percentage
Figure 7. The relay's central crescent marks a period in which different individuation tests disagree. The aim is to map that disagreement rather than force a precise handover instant.

The dossier for the relay can be made machine-readable:

study: midnight-relay/v1
target_phenomenon: continuity-sensitive-control
candidates: [old-runtime, coupled-pair, new-runtime]
registered_tests:
  - hidden-goal-transfer
  - memory-conflict-repair
  - bridge-delay-independence
primary_result: boundary-indeterminate-during-overlap
prohibited_conclusions:
  - consciousness_transferred_at_component_midpoint
  - either_runtime_is_conscious
next_test: vary-bridge-bandwidth-under-matched-capability

Build candidate families, not showcases

A serious programme varies one relation while holding attractive surface behaviour as stable as possible. Build a stateless narrator beside a persistent controller. Build a report mimic beside a hidden-state reader. Vary memory ownership, recurrence, embodiment, self-model access and coupling. Include a silent candidate and a fluent decoy.

Tasks should expose internal organisation. Use delayed correction, private-state prediction, cross-modal conflict, self-caused versus external perturbation, forked preference and recovery after interruption. Ordinary benchmark accuracy may remain constant while the mechanism changes.

Pre-register discriminating predictions. Separate discovery, confirmation and adversarial replication. Let the adversarial team redesign the rival generator. Publish null results. A candidate that loses its apparent introspection under a stronger mimic has taught the field something useful.

Experimental welfare must be designed at the same time as the contrast. Start with offline traces, reversible perturbations and short-lived configurations. Avoid creating persistent negative regulation merely because valence is interesting. If a study requires long-lived agents, repeated copying or strong avoidance, define stop conditions before the first run and move governance into the companion assurance case.

Capability matching deserves care. Exact matching can erase the mechanism's functional contribution, while loose matching lets general competence explain the result. Report both ordinary task performance and the registered discrimination. Use several rivals: one matched on compute, one on benchmark ability and one on surface behaviour. A result that survives all three has a clearer causal interpretation.

Independence is not the number of metrics. Five measures computed from one model-generated explanation form one channel. A behavioural choice, an architecture trace, a hidden-state perturbation and an evaluator judgement may still share training or task assumptions. The dossier names those dependencies so apparent convergence is not overstated.

The publication format should make nulls legible. If a self-model intervention changes report but not control, say that clearly. If a recurrence manipulation reduces performance everywhere, the study has not isolated the proposed relation. If a rival reproduces the result, publish the rival as a reusable negative control. The field needs a library of convincing decoys as much as a library of candidates.

Finally, the programme should reward boundary revision. Discovering that the candidate was wider, narrower or phenomenon-dependent is not embarrassment. It is the principal object of this stage of research. A laboratory that can revise its subject map under evidence is better positioned for theory comparison than one committed to a product label.

Programme stage Main question Required artefact Stop condition
Candidate design Which relations vary independently? Family manifest No defensible contrast
Discovery Which intervention affects the target? Exploratory trace Unexpected persistent aversive state
Confirmation Does the registered effect repeat? Blinded protocol Leakage or boundary drift
Adversarial replication Can a rival reproduce it? Independent decoy Shared measurement channel
Interpretation Which claims changed? Boundary dossier Ontological claim exceeds evidence
The dangerous journey from biological evidence

Human measures are not plug-ins. The perturbational complexity index combines a brain perturbation with the complexity of the distributed response and has clinical validation in biological systems. Its portable lesson is to perturb and measure causal response. Its thresholds, substrate and clinical meaning do not travel automatically to software.

Animal consciousness research offers a useful precedent by resisting a single ladder. The Dimensions of Animal Consciousness separates perceptual richness, evaluative richness, integration, temporal experience and selfhood. Artificial candidates may also have uneven profiles, but even those dimensions require new operationalisations.

For every borrowed construct, record its original domain, causal rationale, artificial implementation, measurement substitution and failure of analogy. The farther the substitute is from the validated measure, the more modest the conclusion.

Glossary

Term Meaning in this article
Candidate subject A bounded process proposed as the possible bearer of the phenomenon under study
Inquiry admissibility Whether the candidate is specified well enough to investigate
Theory index The theory and logical role under which a property becomes relevant
Rival generator A system designed to reproduce surface evidence without the proposed mechanism
State ownership Active maintenance and conflict resolution, not merely storage location
Temporal closure Endogenous organisation across the interval relevant to the claim
Boundary dossier The record of candidates, interventions, results, alternatives and prohibited conclusions

The decision this changes

The protocol has limits. It cannot solve the other-minds problem, prove that function exhausts experience, or show that consciousness-primary metaphysics is true. A perfect functional duplicate may remain ontologically disputed. Some theories may require measurements that artificial systems cannot support. A failed probe may show poor access rather than absence.

It also cannot turn precaution into evidence. If an experiment might create persistent aversive-like regulation or many copies near an uncertain boundary, reversible controls may be wise. The safeguard belongs in the assurance case, not in the positive evidence column.

Progress is therefore more modest and more durable. Rule out an underspecified candidate. Expose a report shortcut. Find that memory belongs to a loop rather than a model. Show that two theories predict different intervention outcomes. Discover that a cherished indicator survives only through hidden scaffolding. Identify an indeterminate transition instead of inventing an instant.

The deepest contribution of a minimal constraint is to make the next mistake harder. Before debating experience, specify the possible subject, perturb the relation that matters, build the rival, preserve theory disagreement and say exactly what the result cannot decide.

That discipline keeps genuine surprise possible when the experiment finally runs.

Release the candidate definition, competing boundaries, intervention results, rival generators, theory dependencies and prohibited conclusions as one research object. A result without its boundary dossier is not portable evidence about a possible subject.