Home · Writing · Consciousness

Thought Experiments as Instruments, not Oracles

A thought experiment earns force by exposing dependencies among premises, bridges and countermodels, not by making one intuition feel decisive.

TLDR

  1. A thought experiment earns force by exposing dependencies among premises, bridges and countermodels, not by making one intuition feel decisive.
  2. A research team builds Room 12. Inside sits Arun, who does not know Chinese. Cards bearing Chinese characters enter through one slot.
  3. A thought experiment gains epistemic force by making a conclusion sensitive to an explicit intervention while keeping rival explanations visible.
  4. Three kinds of variation should therefore be labelled. A causal variation changes a variable within a shared model and predicts downstream effects.
  5. The opening intuition is compelling because first colour experience appears to change Mary. But “she learns something” is not yet a typed claim.
Figure 1. The philosophical wind tunnelConceptual landscape
What this figure changes: treat narrative vividness as the airflow that makes dependencies visible, not as the measuring instrument itself. The output is a map of load-bearing assumptions.Illustrative conceptual figure.
Part I

What the wind tunnel measures

Imagine testing an aircraft wing. The tunnel does not ask whether the wing looks as if it ought to fly. It holds some conditions stable, changes others and records how the system responds. A useful philosophical case can do something similar, but only after its hidden controls are made visible.

A thought experiment is usually presented as a small world. We are told which facts hold, which feature has changed and what an observer should judge. Its narrative form helps us build a mental model and propagate consequences through it. Philosophers disagree about the epistemology of this process. Norton treats many thought experiments as arguments in picturesque form [1]. Nersessian emphasises simulation in a mental model [2]. Stuart argues that they can increase understanding even when their contribution is not exhausted by a compact deduction [3]. The disagreement matters, but it need not be settled before using the instrument well.

Across these accounts, a disciplined case contains at least two separable operations. The scenario contrast tells us what is varied and what is held fixed. The verdict bridge tells us why a reaction to that contrast should support a further conclusion. Most misuse occurs when the first is explicit and the second remains hidden.

Four jobs that should not share one standard

Some thought experiments are consistency probes. They combine claims a theory already accepts and show that the combination produces tension. The output is logical: at least one commitment must be revised. A case of this kind does not need a survey of intuitive reactions, but it does need a faithful reconstruction of the theory.

Others are concept separators. They hold ordinary co-occurrences apart so that two properties can be discussed independently. Mary separates exhaustive physical description from first-person encounter. Teletransportation separates psychological continuity from uniqueness. Such cases can succeed even when the separated combination is not physically buildable. Their secure yield is that our concepts and explanatory demands are not identical, not that nature realises every described separation.

A third group functions as modal arguments. Here the path from coherent description to possibility is itself substantive. A modal case must state the grade of conceivability claimed, explain why hidden contradiction has been excluded and defend the relevant connection between conceivability and possibility. Without this work, “I can imagine it” may report only that the contradiction is not obvious to the imaginer.

A fourth group is best treated as model discriminators. The case places rival explanations under a controlled variation and asks which predicts the changed outcome. This is the closest philosophical analogue to a wind-tunnel test. It is also the most directly useful for science and engineering, because the imagined variation can often become a simulation, intervention or failure-injection experiment.

Prediction: when readers disagree, can you identify whether they are disputing consistency, concept application, modality or causal discrimination? If not, the case is under-typed.

From a reaction to a conditional result

Suppose a case produces the reaction, “The duplicate has no experience.” That reaction may be psychologically strong. It is still only the beginning of analysis. We must ask whether the duplicate was described coherently, whether the description already smuggled in absence of experience, whether conceivability tracks possibility and whether physicalism is being defined as a necessary relation across possible worlds. Each question opens a separate load-bearing span.

Mechanism

A thought experiment gains epistemic force by making a conclusion sensitive to an explicit intervention while keeping rival explanations visible. When the intervention, invariants or bridge change, the conclusion should change for a stated reason. If the verdict survives only because the story keeps repeating it, the case is persuasive theatre rather than a controlled instrument.

Figure 2. The hidden span between case and conclusionCausal cutaway
Practical implication: identify each bridge support before asking whether the conclusion follows. Different objections attack different spans, so “I reject the thought experiment” is too coarse a response.Illustrative causal cutaway.

Not every imagined change is a causal intervention

The wind-tunnel metaphor has a limit. In an engineering tunnel, changing angle of attack while holding air density fixed is an ordinary intervention. Some philosophical cases instead remove a property that may be constituted by what is held fixed. “Keep every physical and functional fact, but subtract experience” is not automatically analogous to disconnecting a wire. It asks whether one complete description necessitates another.

Three kinds of variation should therefore be labelled. A causal variation changes a variable within a shared model and predicts downstream effects. A constitutive variation asks whether a property can be absent while its proposed realising structure remains. A modal recombination asks whether descriptions that travel together in the actual world can come apart in any possible world. The evidential rules differ. Causal variations may become experiments. Constitutive variations require an account of realisation. Modal recombinations require a defended theory of possibility.

Confusing these kinds creates false confidence. A zombie world can pressure a constitutive claim even though no laboratory could switch consciousness off while freezing every physical fact. A teletransporter can expose branching relations even if perfect copying is technologically remote. Their value lies in the dependency being tested, but their distance from an ordinary intervention must remain visible.

Toy worked example: the silent alarm

A machine has a red lamp wired to a smoke sensor. In World A the lamp illuminates because smoke closes the circuit. In World B a technician secretly illuminates the same lamp. The visible output is identical. The case exposes a distinction between indicating smoke and merely displaying the signal associated with smoke. It does not prove that every red lamp lacks representational content. To reach that conclusion one would need a criterion linking content to causal history, use, consumer system or designer intention. The hidden bridge, not the lamp, carries the philosophical burden.

The wind-tunnel view recognises three legitimate yields. A case can expose a distinction that ordinary language hides. It can show that a conclusion follows conditionally from a set of premises. It can also generate a discriminating experiment, model or design requirement. What it cannot do is convert the felt obviousness of an imagined verdict into unearned evidence for a world-level claim.

Part II

Four cases, four hidden bridges

Zombies, Mary, teletransporters and Chinese rooms are often grouped as “intuition pumps”. That label is too blunt. Each case changes a different property, targets a different conclusion and fails under a different kind of countermodel.

1. Zombies: from duplication to possibility

Thought experiment

Imagine a world physically identical to ours, particle for particle and function for function, yet with no phenomenal experience. Its inhabitants speak about pain, avoid injury and publish papers about consciousness, but there is nothing it is like to be them. If such a world is genuinely possible, phenomenal facts do not logically supervene on physical facts.

Chalmers’ modal argument is not simply “zombies seem imaginable, therefore physicalism is false”. In compressed form, it moves from a complete physical truth P, plus the claim that phenomenal truth Q is absent, through conceivability and possibility to a conclusion about metaphysical entailment [4]. The case is powerful because it forces physicalism to say where that route breaks.

The first pressure point is whether P and not-Q has been positively conceived or merely described without an obvious contradiction. The second is whether ideal conceivability entails metaphysical possibility. The third is whether the relevant physicalist thesis requires a priori entailment or only necessary identity that can be known a posteriori. Type-A physicalists deny that the zombie world is ideally conceivable. Type-B physicalists may grant an epistemic gap while denying a metaphysical gap.

The zombie case exposes the price of each position: anti-physicalism needs a defended modal bridge, while physicalism must explain why an apparently coherent epistemic gap does not mark an ontological one. The case does not settle that exchange by asking which world-picture feels easier to imagine.

Figure 3. Where the zombie argument can breakModal pressure bridge
Type-A: deny ideal conceivabilityType-B: deny modal passage
How to use it: do not treat “I can imagine a zombie” as one indivisible premise. Separate descriptive coherence, conceivability grade, modal bridge and target version of physicalism.Argument reconstruction after Chalmers; objection families are schematic.

2. Mary: from learning to ontology

Thought experiment

Mary is a brilliant colour scientist confined to a black-and-white environment. She learns all physical information about colour vision. When she first sees red, does she learn something? Jackson’s original knowledge argument says yes, then infers that not all information is physical information.

The opening intuition is compelling because first colour experience appears to change Mary. But “she learns something” is not yet a typed claim. Does she acquire a new proposition? A recognitional ability? Acquaintance with a quality? A new way of presenting an old physical fact? A memory or imaginative capacity that could not be installed through description alone? These are not verbal evasions. Each predicts a different relation between prior information and later competence.

Jackson’s original formulation explicitly moves from Mary’s post-release learning to the claim that physical information is incomplete [5]. Lewis’ ability hypothesis argues that experience can instead add abilities to remember, imagine and recognise without adding non-physical propositional information [6]. Other physicalist replies appeal to acquaintance or new phenomenal concepts. Jackson later accepted a physicalist response [7], which is itself revealing: the original author’s changed verdict did not change the stipulated room. It changed the bridge from Mary’s post-release gain to an ontological conclusion.

Mary exposes a real explanatory demand: a complete third-person description may leave open why first-person encounter changes what a knower can do, recognise or grasp. Whether that residue is a non-physical fact, a new epistemic relation or a new ability remains the contested step.

Figure 4. Mary’s gain passes through a knowledge prismAnnotated prism
The test it suggests: type Mary’s gain before drawing an ontological conclusion. Only one branch directly yields the original anti-physicalist conclusion; every branch inherits explanatory work.Interpretive synthesis after Jackson and Lewis.

3. Teletransporters: identity, continuity and what matters

Thought experiment

A scanner records every relevant feature of your body and mind, destroys the original and reconstructs an exact counterpart on Mars. The counterpart wakes with your memories, projects and character. Is that survival? Now modify one variable: the scanner no longer destroys the original, and two continuers exist.

The first case tempts a yes-or-no answer about identity. The second reveals why that answer may be the wrong output. Numerical identity is one-to-one: one earlier person cannot be strictly identical to two later people. Psychological continuity, causal lineage, memory, responsibility and rational concern can branch. Parfit uses such cases to ask whether identity is the further fact that matters, or whether less-than-identity relations carry much of the practical importance attributed to it [8].

A pattern theorist may say the Mars continuer is you until branching forces a revision. A bodily-continuity theorist may say the original remains you and the replica never was. A reductionist may say the demand for one deep fact outruns the relations available. The case therefore exposes which criterion is doing the work. It does not let a shiver at the destruction chamber decide metaphysics.

The practical yield appears when identity questions are decomposed. A copying system may preserve memories while breaking legal title. A forked software process may preserve commitments while producing conflicting later preferences. An archive may preserve information while failing to preserve a subject. Teletransportation teaches that “same person” is often an overloaded decision variable hiding lineage, continuity, uniqueness, ownership and welfare.

Figure 5. Branching software spacetime for personsTimeline and fork
Design consequence: when continuity branches, replace the single identity question with typed questions about causal lineage, memory, commitment, title, responsibility and welfare.Conceptual reconstruction after Parfit.

4. Chinese rooms: locating the candidate system

Thought experiment

Searle’s operator follows a program for manipulating Chinese symbols and passes a behavioural test without understanding Chinese. The intended conclusion is that formal symbol manipulation, by itself, is not sufficient for understanding or intentionality.

Searle presents the room to argue that formal program execution is insufficient for intentionality or understanding [9]. The phrase “by itself” carries much of the burden. Which properties count as merely formal? Is the candidate the operator, the operator-plus-rulebook, the entire room, or a room coupled to perception and action? What causal organisation would count as implementing rather than merely describing a program? The systems reply argues that the whole organised system, not the English-speaking operator, is the relevant candidate. Robot and brain-simulator replies alter grounding or causal structure. Searle responds that internalising the system would still leave the person without understanding.

These exchanges are often treated as competing intuitions. They are better read as a boundary and sufficiency audit. Searle asks whether syntax is sufficient for semantics. The systems reply asks whether his introspective access to one component can settle a system-level property. Embodied replies ask whether symbol use acquires content through perception, action and world-involving history. Biological naturalism asks whether the causal powers of the implementing substrate matter.

The Chinese room exposes at least three independent choices: the boundary of the candidate, the implementation relation and the criterion for understanding. Unless those are typed separately, “the room understands” and “the room does not understand” remain verdicts generated by different unseen instruments.

Figure 6. The Chinese room as a moving-boundary cutawaySystem anatomy
Research consequence: declare the candidate boundary before assigning understanding, agency or experience. A component-level introspective absence does not automatically settle a system-level property.Conceptual reconstruction after Searle and major reply families.
CaseControlled contrastHidden bridgeWhat it securely exposesWhat remains open
ZombiesPhysical and functional duplication, phenomenal differenceIdeal conceivability to metaphysical possibilityThe epistemic and modal commitments of physicalismWhether the zombie world is possible
MaryComplete physical information, then first experienceNew epistemic gain to non-physical factDifferent kinds of knowledge and accessOntology of the gain
TeletransporterContinuity with destruction, then branchingCriterion of identity to survival or concernIdentity, continuity and practical concern can divergeWhich relation grounds which norm
Chinese roomCompetence with rule-governed symbol manipulationChosen boundary and sufficiency criterionOutput matching underdetermines mechanism and semanticsWhich organisation constitutes understanding

Read the result at the strength the case earned

The four cases can all be successful without producing their most famous conclusion. A zombie case succeeds as a modal audit when it identifies the exact conceivability principle an anti-physicalist needs and the exact epistemic debt a physicalist inherits. Mary succeeds as a knowledge audit when rival accounts state what changes at release. Teletransportation succeeds as an identity audit when governance rules stop treating memory, lineage and ownership as interchangeable. The Chinese room succeeds as a system audit when a claim about understanding names the candidate and the implementation relation.

This is not a retreat to the harmless statement that “the issue is complicated”. Each audit eliminates bad arguments. It tells us which thesis is genuinely at stake, which premise a critic must attack and which evidence would be relevant. A case becomes weaker only if its original value depended on skipping those steps.

The correct question is therefore not, “Did the thought experiment prove its headline?” Ask instead: “Which rival positions can reproduce every stipulated fact, and at which bridge do their predictions or commitments diverge?” That question preserves the imaginative compression of the case while replacing oracle language with an auditable argument.

Part III

Countermodels turn reactions into arguments

A countermodel is not merely another story with the opposite ending. It preserves enough of the original case to reproduce its observation, then changes the bridge, boundary or mechanism so that the proposed conclusion no longer follows.

Consider Mary. An unhelpful response says, “I intuit that she learns nothing.” A useful countermodel grants the post-release transformation but identifies it as acquiring recognitional and imaginative abilities. It must then explain why those abilities were unavailable in the room and how they arise from experience. Likewise, a response to the Chinese room is not improved by declaring that the room understands. It improves when it specifies the system boundary, semantic relation and causal capacities that the operator alone lacks.

The strongest countermodel matches the surface evidence while reversing the explanatory route. This is why countermodels are constructive. They reveal what additional observation would discriminate the original account from its rival.

Figure 7. One observation, four rival generatorsCountermodel lattice
Boundary condition: when the same observation has several generators, ask which intervention would make their predictions diverge. A matched countermodel blocks an inference even when it does not establish the rival theory.Method proposed here.

A countermodel must pay four bills

Match. Preserve the feature that made the original case appear evidential. A reply to Mary that denies any change after release avoids the central pressure rather than modelling it. A reply to the Chinese room that gives up fluent behaviour no longer explains the observation the room was built to isolate.

Mechanism. State how the rival generates the matched observation. “It is only an ability” is incomplete until the ability account explains why experience creates recognitional or imaginative competence that exhaustive description did not. “The whole system understands” is incomplete until the system account identifies semantic use, grounding, integration or counterfactual sensitivity.

Break condition. Name a change under which the original and rival accounts cease to agree. For Rook-7, scrambling the sensor-to-action relation while preserving fluent output may separate a world-grounded account from a text-only substitute. For teletransportation, branching separates continuity from uniqueness. A countermodel without a break condition may expose underdetermination, but it cannot yet guide inquiry.

Residual debt. Record what the countermodel still owes. The ability hypothesis may preserve physicalism while leaving a question about why physical description cannot confer the relevant ability. A systems reply may block Searle’s component argument while leaving the constitutive theory of understanding unspecified. Blocking an inference is not the same as establishing the rival explanation.

Intuition becomes a measurement object

Experimental philosophy has reported that some judgements about philosophical cases vary with culture, order, framing or participant characteristics [10] [11]. Some early effects have not replicated cleanly [12]. The responsible lesson is neither “intuition is worthless” nor “professional intuition is immune”. It is that the production conditions of a verdict may themselves require study.

For a philosophical wind tunnel, variation in verdicts is diagnostic. If moving one sentence changes most judgements, the sentence may be doing causal work that the theory ignored. If trained readers converge only after an argument is supplied, the argument may carry more weight than the initial case response. If rival descriptions preserve the formal structure but reverse the intuitive verdict, narrative perspective is a variable rather than a transparent window.

Evidence status

Established: the four cases have generated stable families of arguments and objections in the published literature. Contested: what philosophical intuitions are evidence of, and how broadly demographic or framing effects generalise. Method proposed here: the scenario-contrast and verdict-bridge distinction, the wind-tunnel protocol and the fragility classes introduced here.

Figure 8. Illustrative verdict-fragility profilesSynthetic categorical chart
Decision rule: diagnose the case-specific source of fragility rather than assigning a single credibility score. The letters mark argumentative centrality, not empirical frequencies or truth probabilities.Illustrative categorical profile; no measured data.

Worked system example: rook-7

Rook-7 is a configured research system. A language model proposes actions. A persistent state service stores task history. Cameras and microphones supply observations. Tools alter a simulated laboratory. A policy kernel permits or refuses actions. An evidence ledger records inputs, proposals, effects and readbacks.

A Chinese-room-style argument aimed only at the language model may establish that next-token transformation is not, by itself, a full account of Rook-7’s world-directed competence. It cannot establish that the configured system lacks understanding, because the chosen candidate excluded sensors, persistent state, action, policy and feedback. The systems reply also cannot establish understanding merely by enlarging the box. It must specify which relations within that larger boundary constitute content and how interventions would reveal them.

Realistic configured-system example

Run two matched versions. In one, scramble the stable mapping between camera states and tool effects while preserving fluent language. In the other, preserve that causal mapping but replace the language model with a less fluent planner. If world-directed correction survives only in the causally stable version, fluent text is not the sole generator. This result would support a grounding-sensitive account of the configured competence. It still would not establish phenomenal experience.

Why configured AI systems make the boundary problem unavoidable

Modern AI claims often slide between a model and the system built around it. A language model may be stateless between calls while an application supplies persistent state. The model may propose an action while a policy service authorises it. Retrieval may supply current facts that are absent from the weights. Sensors and tools may close a world-directed feedback loop. A thought experiment aimed at “the AI” becomes uninterpretable unless those components and relations are declared.

The wind-tunnel record prevents two opposite mistakes. The first assigns a system-level capacity to the model because the whole application performs well. The second denies a system-level capacity because one component lacks it in isolation. For intelligence and agency, configured-system interventions can often decide which component is causally necessary. For phenomenal experience, the same interventions provide theory-relative evidence only after a candidate boundary and measurement bridge are supplied.

This matters operationally. If the target is responsibility, the relevant boundary may include authority, action contracts and effect receipts. If the target is semantic competence, it may include stable perception-action relations and correction across time. If the target is consciousness, neither fluent report nor enlarged system boundary settles the issue. A larger box is not a stronger argument unless the added relations explain the target property.

Worked failure example: the impossible control

A case says, “Hold every physical and functional fact fixed, remove consciousness, and now observe that nothing functional changes.” This is a legitimate conditional specification only if the joint description is coherent. It cannot also be treated as an experimentally established intervention. When the variable to be removed may be constituted by the properties held fixed, the thought experiment is testing a metaphysical dependency claim, not simulating an ordinary causal manipulation.

Part IV

Run the philosophical wind tunnel

The instrument below turns a case into a typed record. It does not score truth. It reports where a verdict depends on stipulation, contested bridges or surviving countermodels.

The protocol begins with a target claim, because the same story can be used against different theses. It then names the candidate boundary, the held-fixed conditions and the changed variable. Premises receive evidential statuses. The verdict bridge is written as a sentence rather than concealed in “obviously”. Finally, the user supplies the strongest countermodel and asks what observation could make its prediction diverge.

Figure 9. From narrative to an auditable argument recordImplementation workflow
What to inspect: a usable case record ends with both an allowed conclusion and a prohibited inference. When no discriminating test is available, the output should remain conceptual or conditional.Implementation workflow for the executable artefact below.
Executable lab

Argument–countermodel wind tunnel

Select a canonical case or edit every field. The result classifies argumentative fragility, not truth or metaphysical probability.

Load-bearing premises and status

How closely does the countermodel match the original observation?

What the artefact tests

The tool tests whether the argumentative load is visible. A positive result means the case has a declared target, controlled contrast, defended bridge and a countermodel that either fails for a stated reason or predicts something different. That permits a conditional conclusion: given these premises and this candidate boundary, the target claim follows or gains support.

A negative result is also useful. “High fragility” means at least one decisive span is intuition-only, several premises are contested, or a matched countermodel survives. It does not mean the conclusion is false. It means the case, as currently specified, cannot discriminate it from a rival. The correct response is to defend the bridge, refine the scenario or find an intervention.

Worked example: mary through the instrument

Start with the target claim: “Mary acquires knowledge of a non-physical fact.” Set the candidate boundary to Mary-before and Mary-after, not to an abstract database of sentences. Hold fixed the stipulated physical information, reasoning capacity and memory of that information. Change direct acquaintance with colour and the competences it enables.

Now type the bridge. The original route requires more than “Mary changes”. It requires that the change be knowledge, that the knowledge be propositional, that its object be a new fact, and that a fact absent from complete physical information be non-physical. Mark the completeness of physical information as stipulated, Mary’s gain as strongly motivated, and the propositional and ontological typings as contested.

Build the strongest countermodel. Grant that Mary can now recognise, imagine and remember red in a new way. Claim that these are abilities or acquaintance relations realised by physical changes, not knowledge of an additional non-physical proposition. This countermodel matches the central observation, so the wind tunnel reports high argumentative fragility. That is not a verdict for physicalism. It is a precise statement that the original scenario does not discriminate the fact hypothesis from the ability or acquaintance family.

The next move is no longer another vote on what Mary “obviously” learns. Specify post-release tasks. Ask whether Mary could infer every proposition she later asserts while still lacking recognitional competence. Ask whether the ability account predicts dissociations between factual answer, imagery, recognition and memory. The thought experiment has now produced a research programme and a cleaner metaphysical dispute.

Open research hypothesis

Arguments reconstructed with explicit bridge principles and matched countermodels should show less framing sensitivity than bare narrative cases, because readers can locate disagreement in a premise rather than compressing it into a verdict. The hypothesis would be strengthened if independent readers converge on the argument’s dependency structure while continuing to disagree about premise truth. It would be weakened if explicit reconstruction leaves verdict shifts unchanged or merely creates post-hoc rationalisations.

TypeScript · core record
type EvidenceStatus =
  | "empirically-supported"
  | "conceptually-defended"
  | "stipulated"
  | "contested"
  | "intuition-only";

type ThoughtExperimentRecord = {
  id: string;
  targetClaim: string;
  candidateBoundary: string;
  heldFixed: string[];
  changedVariable: string;
  verdictBridge: string;
  premises: { statement: string; status: EvidenceStatus }[];
  strongestCountermodel: string;
  discriminatingTest?: string;
  permittedConclusion: string;
  prohibitedInference: string;
};

// Invariant: never emit a world-level conclusion unless
// the bridge is explicit and surviving countermodels are recorded.

The schema exposes what prose often hides. It intentionally contains no truth score, consciousness score or intuition weight.

Apply the protocol without the interactive form

Write the target claim in one sentence. Draw the candidate boundary. List every condition held fixed, then name the single contrast intended to do work. Translate the intuitive verdict into a bridge principle. For each premise, mark whether it is stipulated, conceptually defended, empirically supported, contested or carried only by intuition. Construct a countermodel that preserves the observation. End with one permitted conclusion, one prohibited inference and one test whose result could weaken your preferred reading.

Figure 10. What conclusion is licensed?Decision instrument
Operational consequence: match the strength of the published conclusion to the bridge and countermodel evidence. Even the strongest quadrant supplies explanatory support under a specified framework, not incorrigible metaphysical access.Practitioner decision matrix.
Decision rule

Publish a thought experiment as an argument packet, not as a verdict: scenario, controlled contrast, candidate boundary, bridge principle, strongest countermodel, discriminating test, permitted conclusion and prohibited inference.

Argumentative conclusion

The decision this changes

A philosopher deciding whether physicalism survives Mary should no longer ask only, “Does she learn something?” A consciousness researcher deciding whether a machine case is probative should no longer ask only, “Can I imagine the system behaving identically without experience?” An engineer deciding whether a fluent model understands should no longer point to a Chinese room without declaring the candidate system. Each decision must move one level down, to the dependency structure that produces the verdict.

This changes how thought experiments should be written, criticised and used in system design. Writers should expose the bridge rather than polishing the intuition. Critics should build a matched countermodel rather than announcing an opposite reaction. Experimentalists should vary narrative cues, system boundaries and causal mechanisms separately. Engineers should translate metaphysical disputes into architecture questions only when a preserved relation and a discriminating intervention have been named.

It also changes the etiquette of disagreement. Two readers can share every observation and still disagree because one rejects a modal bridge, another chooses a wider candidate boundary, or a third types the epistemic gain differently. The argument record lets them locate that disagreement without pretending that one reader failed to “see” the case. This makes null results and persistent disagreement informative: they identify the span that inquiry has not yet secured.

The four cases then become more useful, not less. Zombies reveal where modal necessity enters a theory of consciousness. Mary reveals that “knowing everything” hides several epistemic relations. Teletransporters reveal that continuity, identity and concern can come apart. Chinese rooms reveal that output, implementation, semantics and system boundary are distinct claims. None needs to be discarded because it fails to deliver an oracle’s answer.

The practical decision is simple: never let the intuitive verdict be the last line of a thought experiment. End with the bridge that would make it valid, the countermodel that threatens it and the observation that could change your mind. At that point imagination has stopped impersonating revelation and started functioning as an instrument.

Glossary

Candidate boundary
The component, organism, configured system or wider loop to which a property claim is assigned.
Countermodel
A rival account that preserves the relevant observation or premises while blocking the proposed explanation or conclusion.
Scenario contrast
The controlled difference between the imagined baseline and comparison case.
Verdict bridge
The principle that converts a judgement about the scenario into a conceptual, modal, metaphysical or practical conclusion.
Positive conceivability
Constructing a coherent situation in imagination, rather than merely failing to detect a contradiction in a sentence.
Permitted conclusion
The strongest conclusion licensed by declared premises, bridge and surviving countermodels.

References

Open the source register and extended notes
  1. Methodology
    John D. Norton. “On Thought Experiments: Is There More to the Argument?” Philosophy of Science 71(5), 1139–1151 (2004). Publisher record.
  2. Methodology
    Nancy J. Nersessian. “Thought Experimenting as Mental Modeling.” Croatian Journal of Philosophy 7(2), 125–161 (2007). Record and abstract.
  3. Methodology
    Michael T. Stuart. “How Thought Experiments Increase Understanding.” In The Routledge Companion to Thought Experiments, 526–544 (2018). Author manuscript record.
  4. Primary argument
    David J. Chalmers. “Does Conceivability Entail Possibility?” In Conceivability and Possibility, 145–200 (2002). Oxford record.
  5. Primary argument
    Frank Jackson. “Epiphenomenal Qualia.” The Philosophical Quarterly 32(127), 127–136 (1982). Journal record.
  6. Primary counterargument
    David Lewis. “What Experience Teaches.” Proceedings of the Russellian Society 13, 29–57 (1988). Repository copy.
  7. Position revision
    Frank Jackson. “Mind and Illusion.” Royal Institute of Philosophy Supplement 53, 251–271 (2003). Cambridge record.
  8. Primary argument
    Derek Parfit. Reasons and Persons. Oxford: Clarendon Press (1984), Part Three. Oxford record.
  9. Primary argument
    John R. Searle. “Minds, Brains, and Programs.” Behavioral and Brain Sciences 3(3), 417–457 (1980). Cambridge record.
  10. Empirical intuition research
    Edouard Machery, Ron Mallon, Shaun Nichols and Stephen P. Stich. “Semantics, Cross-Cultural Style.” Cognition 92(3), B1–B12 (2004). Journal record.
  11. Empirical intuition research
    Stacey Swain, Joshua Alexander and Jonathan M. Weinberg. “The Instability of Philosophical Intuitions: Running Hot and Cold on Truetemp.” Philosophy and Phenomenological Research 76(1), 138–155 (2008). Journal record.
  12. Replication boundary
    Adrian Ziółkowski. “The Stability of Philosophical Intuitions: Failed Replications of Swain et al. (2008).” Episteme 18(2), 328–346 (2021; first published online 2019). Cambridge record.