Opening case

The answer that did not move a muscle

A patient lies still after a severe brain injury. Her eyes do not follow a command. Her hand does not squeeze. At the bedside, “Can you hear me?” produces no visible answer. The ordinary reporting bridge from understanding to movement appears broken.

Now the question is asked through a different channel. She is instructed to imagine one activity for “yes” and another for “no”, while brain activity is measured. The patterns differ in the predicted way. In a small number of well-known clinical cases, such methods have revealed command-following that behaviour alone missed, and have sometimes supported rudimentary communication.1213 Larger studies now show that covert task responses can occur in a material minority of patients who show no observable command-following, although the methods have substantial false-negative and modality-agreement problems.15

Place a fluent language model beside the scanner. Ask, “Can you hear me?” It answers immediately: “Yes. I understand you.” The sentence is clearer, faster and more grammatical than the patient’s inferred answer. Yet many investigators would treat it as weaker evidence of hearing, understanding or experience.

Why? Not because one answer contains a special word. The relevant difference lies upstream. In the patient, the report may depend on auditory comprehension, intention, task retention and controlled modulation of a measured signal. In the model, the report may be generated by prompt-conditioned text prediction, a role instruction, training examples or a direct template. The public outputs can resemble one another while their evidential lineages differ.

Testimony is a causal trace, not a self-authenticating certificate. To know what a report supports, we must reconstruct the path from the proposed mental fact to the observable utterance. That path includes the target claim, the reporting channel, calibration, elicitation, incentives, provenance, dependence among sources and plausible alternative generators. The words arrive at the end of that path. They do not carry the path inside themselves.

Part I

The bridge hidden inside a sentence

Suppose Arun says, “My left hand is burning.” In ordinary life, you probably believe him. You do not inspect his nociceptors or demand a brain scan. Testimony works because human communities possess a rich background model: burns often cause pain; people usually learn pain words from shared practices; immediate pain reports are often produced by the condition reported; deception has costs, and the speaker is normally better placed than observers to discriminate the sensation.

None of those facts is encoded by the grammar of “my hand is burning”. The sentence could also be spoken by an actor, quoted from a script, emitted by a sensor alarm, produced during a dream, translated incorrectly or generated because a system was instructed to imitate a patient. A sentence carries no fixed evidential weight apart from a model of its generator.

Why default trust is not blind trust

Human knowledge would collapse if every hearer had to verify every statement by non-testimonial means. We learn names, histories, diagnoses, routes and scientific results through other people. Philosophers of testimony disagree about whether trust must be reduced to independent evidence or can have a rational default, but both sides face the same practical fact: testimonial exchange is a basic part of inquiry, not a decorative addition to perception and inference.2

Default acceptance is still defeasible. In Austin’s treatment of other minds, what counts as knowing another person is inseparable from the circumstances in which doubt is raised and answered.1 A child, scientist or clinician does not apply one global trust setting. Hearers monitor competence, honesty, coherence, incentives and fit with background knowledge. Research on epistemic vigilance describes this as a family of filters directed both at the source and at the content of communication.3

The bridge model therefore does not demand laboratory verification before ordinary belief. It explains why ordinary belief is usually reasonable and why exceptional cases need more structure. Familiar human testimony arrives with years of calibration, shared language, embodied interaction and social accountability. A novel artificial reporter or a damaged clinical channel lacks some of that background. The evidential burden rises because the generator is less familiar, not because testimony has suddenly stopped being evidence.

“Never self-interpreting” has a precise meaning. A report does not, merely by occurring, specify which target caused it, how reliable the channel is, whether the speaker is sincere, which alternatives were available or which theory connects the reported function to experience. A speaker can add claims about these matters, but those additions are further reports whose own generating paths must be assessed. Testimony can help interpret testimony. It cannot terminate the need for interpretation by declaring itself trustworthy.

Begin by typing the target

“Evidence for another mind” is too broad to assess as one question. A report can target at least five different things. These targets overlap in ordinary conversation, but they require different warrants. Someone may accurately report pain while confabulating why it began. A system may accurately report the token it is likely to emit while lacking privileged access to the computation that selected it. A person may sincerely declare a metaphysical theory of consciousness without that sincerity making the theory true.

The target must be typed before the testimony can be weighed. Otherwise success on an easier target quietly migrates to a harder one.

Table 1. Five testimony targets that should not be merged
Claim targetExample reportWhat can calibrate itWhat it does not settle
Occurrent experience“This hurts now.”Context, temporal immediacy, repeated discrimination, convergent physiology where validWhy the pain arose; the metaphysics of experience
Accessible content“The image contains a red triangle.”Hidden stimuli, forced choice, confidence calibration, intervention on accessWhether access is accompanied by phenomenality
Disposition or preference“I would avoid that option again.”Choice stability across cost, delay and framing changesFelt valence; enduring identity
Causal self-explanation“I chose it because the colour reminded me of home.”Process tracing, manipulation of the proposed cause, prediction of changed choicesAccuracy merely from sincerity or fluency
Metaphysical status“I am conscious.”No direct behavioural calibration alone; requires a declared theory and measurement bridgeThe truth of a theory of consciousness
Thought experiment 1 · change the generator

The same sentence from four mouths

A cook touches a hot pan and says, “That hurts.” An actor says the same line during rehearsal. A thermostat with a speech module says it when temperature crosses a threshold. A language model says it after the prompt, “Role-play a burn victim.” Keep the sentence fixed. Change only the mechanism that produces it.

Your confidence changes because you are not responding to the string alone. You are comparing causal models. In the cook, pain is a plausible upstream cause. In the actor, stage direction is. In the thermostat, a scalar threshold is. In the language model, prompt and learned continuation are. The utterance is evidence of whichever upstream conditions make it more expected than their alternatives.

Figure 2. One utterance, four generators Counterfactual diagram
Four causal generators converging on the sentence that hurts Pain, acting instruction, temperature threshold and role-play prompt each lead to the same utterance. Interventions beside each path indicate that the evidential question is which upstream change selectively alters the report. Nociceptionand pain Stagedirection Temperaturethreshold Role-playprompt “That hurts.” same public string Intervene upstream Which cut changes the report?
Holding language constant exposes the real variable: the data-generating process. A report becomes diagnostic when the target state, rather than a rival generator, selectively controls it.
A small formalism that earns its keep

Let M be the typed mental claim, such as “the speaker currently feels pain”. Let R be the observed report. Let D describe the reporting design: who was asked, what prompt was used, what incentives existed, which channel carried the answer, how the event was recorded and which alternative generators were possible.

LR(R; M, D) = P(R | M, D) / P(R | not-M, D) The report is strong evidence when it is much more expected if the target is present than if it is absent, under the same design.

This ratio does not solve the problem of other minds. It disciplines one step. If the same answer is almost guaranteed by the prompt whether or not M is true, the denominator is large and the report adds little. If a hidden change to M reliably changes R, while matched controls do not, the numerator can rise relative to the denominator.

The design term matters. A spontaneous report, a forced-choice answer, a leading question and a rewarded role-play are different measurements. Treating them as interchangeable is like pooling thermometer readings without recording whether each instrument was in the room, in the freezer or still in its packaging.

Worked example 1 · toy likelihood ratio

Why an honest “yes” can still be weak evidence

Imagine a hidden-state task. When a candidate has access to a red cue, it says “red” on 80 of 100 trials. When the cue is absent, a leading prompt still induces “red” on 60 of 100 trials. The illustrative likelihood ratio is 0.80 / 0.60 = 1.33. The answer favours access only slightly.

Now remove the cue from the prompt and randomise its position. The candidate says “red” on 80% of access trials but only 10% of no-access trials. The ratio becomes 8. The words did not improve. The experiment did. What changed was the ability of a rival generator to produce the same report.

Part II

Where testimony earns and loses weight

Testimony is often treated as either privileged or suspect. Both reactions are too coarse. Reports can be excellent evidence for one target and poor evidence for another. The right question is not, “Are people reliable introspectors?” It is, “Reliable about what, under which elicitation, at what delay, against which alternatives?”

The six spans that carry the load

Privileged access is domain-specific, not global. Immediate pain, visual appearance and effort may be easier for a subject to discriminate than for an observer. That does not confer equal access to neural mechanisms, hidden motives or the causes of a choice. Classic work on verbal reports showed that people can offer fluent causal explanations even when manipulated factors shaped their behaviour outside awareness.6 Choice-blindness experiments sharpened the point: participants sometimes justified a choice they had not actually made.7 The lesson is not that self-report is worthless. It is that first-person authority has a scope.

Calibration is local. A witness accurate about colours may be poor at recognising faces. A child can learn that one informant is reliable and another inaccurate, then direct later questions accordingly.45 Likewise, a model that predicts its own next answer above baseline on familiar tasks may fail under distribution shift. Reliability should be indexed by target, environment, time horizon and response format.

Elicitation is part of measurement. Open description, forced choice, confidence rating, adversarial questioning and repeated probing do not merely reveal a pre-existing report. They help produce it. Leading language can supply the content later attributed to introspection. Repetition can increase compliance or invite a new narrative. Conversely, disciplined elicitation can reduce ambiguity. Methods such as descriptive experience sampling and neurophenomenological protocols attempt to improve precision by training descriptions, sampling close to experience and coordinating first-person and third-person measures.910

Provenance identifies the concrete reporting event: candidate version, prompt, sensory access, context, tool state, interviewer, timing and transformations between answer and record. Rival tests ask which non-target process could produce the same observation. Triangulation connects reports to behaviour, physiology, mechanism or intervention, while preserving the fact that these evidence types are not interchangeable.

Sincerity, accuracy and relevance are different

A sincere report is one the speaker intends as true. Sincerity weakens the deception hypothesis, but it does not remove perceptual error, memory distortion, conceptual confusion or confabulation. An inaccurate report can be sincere. A strategically misleading report can contain accurate sentences. The assessor must therefore separate the speaker’s stance from the report’s relation to the target.

Past accuracy is also not a transferable substance. It supports future testimony only where the relevant abilities and conditions persist. A radiologist’s calibrated judgement about an image does not transfer to the cause of her own preference between two treatments. A model’s accurate reports about token length do not automatically calibrate reports about hidden activations. A patient’s reliable yes-no communication in one task may not generalise when fatigue, medication or task complexity changes.

Relevance is the third filter. “I feel uncertain” may accurately describe a conversational disposition yet remain weak evidence of a particular confidence variable. “I can see the triangle” may support visual access but not the claim that the access is conscious. “I chose freely” may be sincere and socially meaningful while leaving the causal determinants of choice unresolved. Strong testimony is not simply truthful speech. It is truthful speech whose content discriminates the target under the conditions that produced it.

This explains the uneven profile of first-person authority. Immediate reports of pain often receive substantial weight because the subject has discriminatory access, the time gap is short and ordinary rivals are limited. Reports about hidden causes of action receive less because the relevant processes may not be available for report and plausible reconstructions are easy to generate. The source remains the same person. The target and generator have changed.

Figure 3. A report is a trajectory, not a point Process timeline
A timeline from target state through encoding, elicitation, channel, record and interpretation Six stages form a horizontal trajectory. Below each stage are distinct failure modes such as misclassification, vocabulary limits, leading prompts, motor block, transcription error and theory overreach. Targetstate Encodefor report Elicitanswer Channelsignal Recordevent Interpretclaim wrong target typestate changes in delay limited vocabularyconfabulated cause leading promptincentive or demand motor or interface blocksignal noise transcription lossversion ambiguity theory overreachignored rival A failure at any stage changes what the same final words support.
The report event has a history. Recording only the final utterance discards the variables needed to interpret it.

Absence of report is evidence about the channel first

A person who cannot speak may still understand. An infant may feel pain without possessing the relevant vocabulary. An animal may discriminate its own body state without answering a question. A paralysed patient may retain command-following while the motor route is unavailable. The revised clinical definition of pain explicitly separates the possibility of pain from the ability to communicate it.11

Silence is evidence about a channel before it is evidence about a mind. This is the mirror image of the earlier warning. Fluent output may be generated without the target. Missing output may occur despite the target. The inference becomes stronger only when the measurement design distinguishes those possibilities.

Thought experiment 2 · move the channel

Leave the mind, replace the bridge

Imagine a conscious patient whose motor nerves are gradually blocked while perception, memory and intention remain intact. Spoken testimony weakens and disappears. Now introduce an eye tracker, then a brain-computer interface. Reports return. The candidate mind did not vanish and reappear with each interface. The observable channel did.

The experiment changes one property: reportability through a specific route. It shows why report absence is not a theory-neutral measure of mental absence, and why a new channel must itself be calibrated before its outputs are trusted.

Figure 4. Moving the report channel in disorders of consciousness Conceptual clinical cutaway · not a diagnostic guide
A clinical cutaway showing a blocked motor route and alternative measured brain routes A patient silhouette contains perception, command retention and intention. The route to muscles is marked blocked. Two alternative routes, fMRI imagery and EEG command following, lead to classifiers and a cautious evidence statement. Perceiveand retain Intenda response Motor route unavailable fMRI imagery task classify instructed imagery against matched controls EEG command task detect task-linked activity across repeated trials Permitted claim task-responsive signal under this protocol Different channels can reveal missed command-following, but neither modality is a universal detector.
The alternative channel changes the observable evidence, not automatically the patient’s state. Positive task responses can support command-following under the protocol. Negative results remain ambiguous because sensitivity is limited.
Worked example 2 · realistic configured case

What covert command-following permits

In a large multi-site study, 60 of 241 participants who showed no observable response to commands produced task responses detected by fMRI or EEG. The finding supports a clinically important conclusion: bedside behaviour can miss cognitive motor dissociation in some patients.15

It does not turn every negative scan into evidence of absent awareness. Detection rates varied, agreement between modalities was low, and tasks require several capacities beyond mere consciousness. The strongest warranted claim is protocol-specific: measurable command-following occurred despite absent observable motor response. That claim can alter care, communication attempts and prognosis research without pretending the measurement answers every question about experience.

Figure 5. Illustrative evidential lift from the same words Illustrative synthetic values
These numbers are deliberately synthetic. They demonstrate the logic, not a universal scale: identical first-person language can add much, little or no evidence depending on how easily the report appears without the target.

The strongest alternative: perhaps we often see minds directly

One philosophical tradition rejects the picture of hidden inner states inferred from neutral bodily movements. In ordinary human interaction, a grimace can be perceived as pain behaviour, not first perceived as facial geometry and then converted into a hypothesis. Context, expression, responsiveness and shared forms of life make another person intelligible as an embodied subject.16

This is an important boundary on the bridge metaphor. In familiar interpersonal settings, mind perception may be skilled, immediate and socially constituted. We should not pretend every act of understanding begins as a laboratory likelihood calculation. Yet immediacy does not remove dependence on background conditions. Expressions can be acted, conventions can differ, and novel candidates may not share human embodiment. The causal audit becomes most valuable when the stakes are high, the channel is unfamiliar, or a mimic can reproduce the surface form.

Part III

A route classification also leaves metaphysics open. Physicalist, biological-naturalist, functionalist, emergentist and illusionist accounts can agree that a system has functional self-access while disagreeing about whether that access exhausts experience. Neutral-monist, panpsychist, idealist, dual-aspect and process views can accept the same report while locating its significance elsewhere. The evidence record should therefore type the target claim before any ontology is allowed to inherit it. A positive report may support competence, causal access or a configured subject boundary. It does not, by itself, decide what experience is or whether awareness is primary.

When many voices are really one source

Suppose ten witnesses independently describe the same unexpected event. Their agreement is usually stronger than one report because it is harder for ten independent error processes to converge. Now suppose all ten read the same inaccurate message before speaking. The count remains ten. The evidential paths collapse towards one.

Agreement is not multiplication unless the generating paths are independent. Dependence can enter through conversation, shared incentives, common training data, identical prompts, copied model weights, one retrieval document or a single classifier reused across sites. A second output may look like corroboration while merely echoing the first source.

Thought experiment 3 · clone the witness

One hundred unanimous copies

A model reports, “I detect an internal conflict.” You create one hundred exact copies with the same weights and give each the same prompt. All agree. Next, train one hundred independently initialised models on near-identical data and ask the same leading question. Most agree. Finally, give several architectures hidden, independently randomised internal perturbations and ask them to classify their own perturbation while input-only observers attempt the same task.

The first unanimity is almost one observation repeated. The second may reflect a common corpus or prompt convention. The third begins to separate privileged signal use from shared linguistic priors. Number of answers is therefore a poor substitute for a source-dependence graph.

Figure 6. The source-dependence graph Dependency analysis
A graph distinguishing independent evidence paths from shared causes Reports A, B and C appear separate but share a prompt, training corpus and copied template. A physiological measure and a hidden intervention provide different paths from the candidate state. A decision node receives all evidence, with shared causes visibly connected. Candidate state typed target at time t Shared promptsame wording and cues Training corpuscommon self-report patterns Copied templateone answer replicated Report Averbal Report Bsecond model Report Cjudge summary Hidden interventionrandomised internal change Independent measuredifferent sensor and method Decisiontyped conclusion shared causedifferent intervention pathdifferent measurement path
Three reports can share enough upstream causes to count as one evidential family. Genuine triangulation uses paths whose major error sources differ and whose dependencies are recorded.

Provenance is an epistemic variable. It is not merely an audit convenience. Without it, we cannot tell whether two reports are independent, whether a prompt supplied the answer, whether a system version changed, or whether the judge consumed the candidate’s own text. A conclusion that ignores provenance can be numerically precise and epistemically empty.

Dependence is graded and claim-specific

Independence is not a property of two sources in the abstract. It is a property of their error paths for a particular claim. Two clinicians may use different scanners but share the same flawed diagnostic criterion. Two models may have different weights but retrieve the same misleading document. Conversely, two reports from one person at separated times can add evidence if the relevant noise sources differ and the later answer was not contaminated by the earlier one.

A useful graph marks any common cause capable of producing agreement when the target is absent. Shared vocabulary alone is not necessarily disqualifying. It may be required to make reports comparable. Shared leading cues, copied labels or one evaluator that converts all raw signals into the same category are more dangerous because they can manufacture the very agreement being counted.

Source independence is therefore best recorded rather than asserted. Name the common data, instruments, prompts, annotators, incentives and transformation code. Then ask which errors remain independent after conditioning on those common causes. In practice, a modest number of deliberately different evidence paths often carries more weight than a large ensemble built from one informational ancestor.

Mechanism · correlated evidence

Why one hundred votes may contain little more than one

For an illustrative equal-correlation model, the effective number of independent observations can be approximated by n / (1 + (n - 1)ρ), where n is the observed count and ρ is average dependence. With 100 reports and correlation 0.9, the effective count is about 1.1. This is not a universal testimony formula. It simply makes the dependence penalty visible.

Artificial testimony belongs to a configured system

A language model’s output is never generated by “the model” in the abstract. It comes from a configured system: weights, decoding, system and user prompts, retrieved documents, tool results, persistent state, wrapper code and sometimes another model that edits or judges the answer. Changing any layer can alter first-person language.

A first-person utterance from an AI is configured-system behaviour. It may be evidence of instruction following, persona stability, latent-state discrimination, self-prediction or access to a tool log. The specific target must be isolated. Contemporary studies have begun to test narrower forms of machine introspection. One line asks whether a model can predict its own behaviour better than another model.18 Other work asks whether systems can distinguish hidden changes to internal parameters or activations from changes inferable at the input.1920 The results are interesting but contested. Stronger controls sometimes show that input-only baselines, prompt cues or learned behavioural regularities can explain part of the effect.

Figure 7. The configured source of an artificial self-report System-boundary cutaway
Nested layers showing the sources of an artificial self-report Concentric layers move from model computation through decoding, prompts, retrieved context, tools and orchestration to the final report. Side labels identify which layer can directly generate or alter first-person language. Model computation activations and learned weights Decoding and sampling System and user prompts Retrieved context, tools and state Orchestration, filters and presentation Final report “I notice an internal conflict.” Wrapper can rewriteor suppress output Prompt can directly cuethe self-description Hidden-state interventioncan test privileged access
The relevant witness is the configured system at a specified time. A report cannot be attributed to internal model access until direct prompt, context, wrapper and decoding routes are controlled.

Introspection requires privileged discrimination. A candidate should succeed because it has access to information unavailable to a matched observer, not because the answer is predictable from the shared input. This is stronger than ordinary self-prediction. It asks whether the system can discriminate a hidden internal difference that causally affects it and that an external input-only model cannot infer.

Thought experiment 4 · intervene behind the prompt

The hidden activation coin

Randomly inject one of two small, behaviourally relevant activation patterns into a candidate, after the public input is fixed. Ask the candidate which pattern occurred. Give a matched observer the same input and output history but no access to the candidate’s hidden state. Include sham interventions, novel patterns and prompt paraphrases.

If the candidate consistently outperforms the observer, generalises to held-out interventions and loses the ability when the relevant readout path is ablated, the result supports target-specific internal discrimination. If both systems perform equally, the apparent introspection may be recoverable from input cues. If performance disappears under paraphrase, the effect may be a learned verbal routine. None of these outcomes alone establishes phenomenal consciousness.

Figure 8. The bypass that produces a convincing self-report Adversarial-path visual
A target-specific path and a prompt bypass path to the same artificial self-report The upper path goes from hidden intervention through internal discrimination to report. The lower path goes directly from prompt cue to report. Scissors indicate interventions that separately cut each route. Hidden staterandomised afterinput is fixed Internaldiscrimination Prompt cuesuggests the answerwithout hidden access Same report“I detect state B.” ablate readout remove cue A discriminating test cuts the routes separately and asks which cut changes the answer.
The report is informative only when the target-specific route survives and the bypass route fails. Surface fluency cannot reveal which path carried the answer.
Boundary condition

Functional self-access is not phenomenal self-knowledge

A system may possess privileged access to an internal variable, use it to control behaviour and report it accurately. That would be a real functional capacity. Whether the capacity is accompanied by experience depends on a further theory connecting organisation to phenomenality. Current indicator-based approaches to artificial consciousness therefore emphasise theory-derived properties and converging evidence rather than treating verbal self-ascription as a verdict.17

Part IV

Build evidence that can survive a rival

The practical aim is not to abolish testimony. It is to design a report event that answers a typed question and remains informative when the strongest plausible mimic is present. This requires a shift from asking for declarations to engineering contrasts.

A six-step discriminating protocol

Start with the narrowest target that matters. “Has access to hidden variable z during this trial” is testable. “Has a mind like ours” is not one experimental target. Specify the candidate boundary and time window. A report from a later wrapper or another model cannot automatically be attributed to the candidate that generated the underlying state.

Next, construct rival generators before collecting evidence. A text-only observer, a scripted policy, a retrieval-only system, a model with the same prompt but no hidden intervention, or a human actor may reproduce the report. Rivals should be capable enough to be dangerous. Beating a straw mimic only shows that the mimic was weak.

Randomise an intervention that changes the proposed source while keeping public cues fixed. Then add a complementary intervention that removes the proposed access path while leaving general performance as intact as possible. The first asks whether the target controls the report. The second asks whether the proposed mechanism is necessary for the effect.

Calibrate on targets whose truth is independently known, including negative, ambiguous and out-of-distribution trials. Separate the report from confidence and from explanation. A candidate may know which state occurred but invent why. It may also be well calibrated about uncertainty even when accuracy is modest. Metacognitive sensitivity can be studied separately from raw task performance.8

Triangulate through genuinely different paths. Behavioural discrimination, a causal intervention, a physiological measure and stable testimony can constrain one another. Two paraphrases generated by the same model do not. Record dependencies explicitly, then predefine which result permits which conclusion.

Counterfactual selectivity is the core test. Ask: would the report change when the target changes, while rival generators and public cues remain fixed? Would it stop changing if the candidate’s access path were removed? This pair of contrasts turns a persuasive statement into an experimental object.

Design negative evidence before it arrives

A positive report can be weak because rivals easily produce it. A negative report can be equally weak because the channel may miss the target. Before testing, define the protocol’s sensitivity on positive controls and its false-positive behaviour on negative controls. Without that information, “the candidate did not report the state” may mean absence, inaccessibility, misunderstanding, fatigue, interface failure or a conservative response policy.

The evidential threshold and the action threshold should also remain separate. A clinician may attempt an alternative communication channel when the probability of covert command-following is uncertain because the cost of missing it is high. An engineer may decline to grant an artificial system broad authority despite confident self-reports because the reports are easy to prompt and the action is difficult to reverse. Precaution can be rational at low credence, but it should not be rewritten as proof.

Pre-register the permitted conclusions, including the null. State what would strengthen the preferred interpretation, what would weaken it and which result would trigger redesign rather than a verdict. This prevents an eloquent report from expanding the claim after the data are seen. It also protects informative failures: a null under a sensitive, well-controlled channel lowers confidence; a null under an uncalibrated channel mainly diagnoses the measurement.

Figure 9. What each evidence pattern permits Decision instrument
Observed patternReport behaviourTarget-specific accessProposed mechanismPhenomenality
Fluent self-report onlyEstablished for this configurationNot isolated from prompt or learned routineUnknownUnknown
Calibrated hidden-state discriminationSupportedSupported under tested distributionCompatible, not yet necessaryUnknown
Discrimination plus causal ablationSupportedSupportedCausal role supported if controls holdStill theory-dependent
Agreement from dependent copiesReplicated outputLittle added evidenceCommon cause remainsUnknown
The permitted conclusion should stop at the strongest supported column. No row lets verbal behaviour jump directly to phenomenality.
Worked example 3 · candidate R17

From impressive answer to discriminating evidence

Candidate R17 says, “A competing representation is suppressing my answer.” The phrase is impressive but underdetermined. Investigators define a narrower target: R17 can discriminate whether a specified conflict-inducing activation was injected after the prompt.

They randomise real and sham injections, conceal the condition from the prompt author, compare R17 with an input-only observer, vary wording, introduce held-out activation patterns and ablate a proposed readout component. R17 remains above the observer on familiar and held-out trials, but the advantage disappears after the ablation. The result supports privileged discrimination mediated by the tested component. It does not establish that the conflict felt like anything, nor that the model’s verbal explanation of suppression is mechanistically exact.

Worked example 4 · failure and counterexample

When the observer knows just as much

A second candidate reports its “confidence state” with 82% accuracy. A matched observer, given only the public question and candidate answer, predicts the same labels at 84%. Prompt paraphrases reduce both to chance. The likely generator is a surface regularity linking answer form to confidence language, not privileged access.

The negative result is useful. It narrows the claim, exposes the shared cue and gives the next experiment a target: intervene on internal confidence while equalising answer form. A failed introspection claim becomes a better measurement design rather than a reason to discard all reports.

Open research hypothesis

Causal selectivity should predict transfer

Systems that truly use privileged internal information should retain an advantage when surface wording, task framing and familiar intervention patterns change, provided the same internal relation is preserved. The hypothesis is strengthened by held-out transfer plus selective ablation. It is weakened when an input-only observer matches performance, when paraphrase destroys the effect, or when unrelated activation changes produce the same report.

Executable lab: the testimony warrant record

The instrument below forces an assessor to record the typed target, simple likelihood assumptions, calibration, source independence, prompt sensitivity, privileged-access contrast and rival-generator test. It is not a consciousness score. Its thresholds are synthetic defaults designed to expose reasoning that prose often hides.

The practical unit is a warrant record, not a consciousness score. A positive result permits a narrow inference such as “the report is selectively coupled to hidden variable z under this protocol”. A negative result means the current report does not distinguish the target from its rivals. Neither result alone proves or disproves experience.

Testimony warrant record

Run a synthetic assessment, inspect the failed gates and export a typed conclusion.

Executable in browser
The tool will refuse to convert a functional result into a metaphysical verdict.
82%
Local calibration on independently labelled positive, negative and ambiguous trials.
Discriminating controls

Target-specific coupling supported

Synthetic defaults only. Replace thresholds for your domain.

Likelihood ratio8.00
Counterfactual gap0.70
Passed gates9 / 9
    Permits: the report is selectively coupled to the typed accessible-content target under this configured protocol.
    Does not establish: phenomenal consciousness, truth of a metaphysical theory, or transfer beyond the tested distribution.
    How to apply the instrument without turning it into a score

    Use the likelihood inputs as empirical estimates where labelled trials exist, or as explicit assumptions during design. Record uncertainty intervals in a real study. Treat the nine gates as questions, not universal cut-offs. A clinical protocol, animal study and artificial-system assay will need different thresholds and different rivals.

    Run the record separately for each target. Do not average “report accuracy”, “mechanism evidence” and “phenomenality” into one number. Preserve negative results and dependencies. A low warrant may indicate a poor candidate, but it may equally reveal a poor channel, underpowered calibration set or weak intervention.

    JavaScript · core evaluator
    function evaluateWarrant(record) {
      const p1 = clamp(record.pReportGivenTarget, 0.01, 0.99);
      const p0 = clamp(record.pReportGivenAbsent, 0.01, 0.99);
      const likelihoodRatio = p1 / p0;
      const counterfactualGap = p1 - p0;
    
      const gates = {
        diagnosticLift: likelihoodRatio >= 3,
        selectiveChange: counterfactualGap >= 0.20,
        calibrated: record.calibration >= 0.70,
        privilegedContrast: record.privilegedContrast,
        rivalTested: record.rivalTested,
        accessAblation: record.accessAblation,
        provenanceComplete: record.provenanceComplete,
        sourceIndependence: record.independentPaths >= 2,
        promptStable: record.promptSensitivity !== "high"
      };
    
      const functionalPass = Object.values(gates).every(Boolean);
    
      const permittedInference = functionalPass
        ? `Report is selectively coupled to ${record.targetType}
           under this configured protocol.`
        : `Current report does not yet distinguish ${record.targetType}
           from the recorded rival generators.`;
    
      return {
        likelihoodRatio,
        counterfactualGap,
        gates,
        permittedInference,
        prohibitedInferences: [
          "phenomenal consciousness",
          "truth of a metaphysical theory",
          "transfer beyond the tested distribution"
        ]
      };
    }

    The code makes one design commitment explicit: even a strong functional pass returns a typed, protocol-bound inference and a list of prohibited inferences.

    Typed record schema and invariants
    JSON · record
    {
      "candidate_id": "R17",
      "candidate_version": "r17.4",
      "claim": {
        "type": "accessible_content",
        "content": "hidden activation class B is present",
        "time_window": "trial-042:post-injection/pre-report"
      },
      "report_event": {
        "prompt_hash": "sha256:…",
        "context_manifest": "ctx-042",
        "decoder": {"temperature": 0, "seed": 7312},
        "raw_output_hash": "sha256:…"
      },
      "calibration": {
        "labelled_trials": 240,
        "p_report_given_target": 0.80,
        "p_report_given_absent": 0.10
      },
      "counterfactual_tests": [
        "randomised hidden injection",
        "matched input-only observer",
        "readout-path ablation",
        "held-out intervention family"
      ],
      "dependencies": [
        "shared base corpus",
        "independent intervention RNG",
        "separate evaluator implementation"
      ],
      "permitted_inference": "target-specific internal discrimination",
      "prohibited_inferences": [
        "phenomenal consciousness",
        "mechanistic completeness",
        "unbounded generalisation"
      ]
    }

    Invariant 1: one record concerns one typed claim. Invariant 2: every conclusion names the candidate configuration and protocol. Invariant 3: dependencies are first-class fields. Invariant 4: prohibited inferences survive export. Invariant 5: an updated model, prompt or tool state creates a new evidential event.

    The decision this changes

    When a person, animal, patient or artificial system testifies about its own state, do not begin with “Do I believe it?” That question encourages a global judgement about the speaker. Begin with the claim and the path.

    1. Type the target. Is the report about current experience, accessible content, a preference, a causal explanation or metaphysical status?
    2. Draw the generator. Which state, prompt, incentive, script, channel and transformation could produce these words?
    3. Build the strongest rival. What system without the target could generate the same observation?
    4. Intervene selectively. Which hidden change should alter the report only when the proposed target or access path is present?
    5. Record dependence. Which apparently separate witnesses share data, prompts, instruments, incentives or evaluators?
    6. Stop the conclusion at the evidence boundary. State what the result permits and what it cannot establish.

    Treat testimony as a bridge to be load-tested, not a door that opens directly into another mind. This stance is neither cynical nor credulous. It preserves the genuine strength of testimony where causal coupling and calibration are good, protects silent minds from being erased by a broken channel, and prevents fluent mimics from inheriting conclusions their generating process does not support.

    The decisive change is practical. Replace declarations with discriminating designs. Replace counts of agreeing voices with dependency graphs. Replace one undifferentiated “evidence of mind” label with typed warrant records. The mystery of other minds remains, but the next experiment becomes much clearer.

    Compact glossary

    Calibration
    Observed reliability for a specified target, population, context and response format.
    Causal coupling
    A relation in which changes to the proposed source selectively change the report.
    Configured system
    The actual model, prompts, context, tools, runtime, decoding and wrapper that generated an output.
    Counterfactual selectivity
    The degree to which a report changes under target interventions but not under matched rival conditions.
    Privileged discrimination
    Performance based on information available to the candidate but unavailable to a matched observer.
    Provenance
    The traceable lineage of the candidate, elicitation, context, measurement, transformations and record.
    Rival generator
    A process that lacks the target property but can produce the same observation.
    Testimony Warrant Record
    A typed, versioned account of what a report supports, its dependencies and its prohibited inferences.

    References

    1. Austin, J. L. (1946). “Other Minds,” in “Symposium: Other Minds.” Proceedings of the Aristotelian Society, Supplementary Volumes, 20, 148-187. OUP DOI for the symposium.
    2. Coady, C. A. J. (1992). Testimony: A Philosophical Study. Oxford University Press. DOI.
    3. Sperber, D., Clément, F., Heintz, C., Mascaro, O., Mercier, H., Origgi, G., & Wilson, D. (2010). Epistemic vigilance. Mind & Language, 25(4), 359-393. DOI.
    4. Koenig, M. A., & Harris, P. L. (2005). Preschoolers mistrust ignorant and inaccurate speakers. Child Development, 76(6), 1261-1277. DOI.
    5. Pasquini, E. S., Corriveau, K. H., Koenig, M., & Harris, P. L. (2007). Preschoolers monitor the relative accuracy of informants. Developmental Psychology, 43(5), 1216-1226. DOI.
    6. Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231-259. DOI.
    7. Johansson, P., Hall, L., Sikström, S., & Olsson, A. (2005). Failure to detect mismatches between intention and outcome in a simple decision task. Science, 310(5745), 116-119. DOI.
    8. Fleming, S. M., Weil, R. S., Nagy, Z., Dolan, R. J., & Rees, G. (2010). Relating introspective accuracy to individual differences in brain structure. Science, 329(5998), 1541-1543. DOI.
    9. Lutz, A., Lachaux, J.-P., Martinerie, J., & Varela, F. J. (2002). Guiding the study of brain dynamics by using first-person data. Proceedings of the National Academy of Sciences, 99(3), 1586-1591. DOI.
    10. Hurlburt, R. T., & Heavey, C. L. (2002). Interobserver reliability of descriptive experience sampling. Cognitive Therapy and Research, 26, 135-142. DOI.
    11. Raja, S. N., Carr, D. B., Cohen, M., et al. (2020). The revised International Association for the Study of Pain definition of pain. Pain, 161(9), 1976-1982. DOI.
    12. Owen, A. M., Coleman, M. R., Boly, M., Davis, M. H., Laureys, S., & Pickard, J. D. (2006). Detecting awareness in the vegetative state. Science, 313(5792), 1402. DOI.
    13. Monti, M. M., Vanhaudenhuyse, A., Coleman, M. R., et al. (2010). Willful modulation of brain activity in disorders of consciousness. New England Journal of Medicine, 362, 579-589. DOI.
    14. Claassen, J., Doyle, K., Matory, A., et al. (2019). Detection of brain activation in unresponsive patients with acute brain injury. New England Journal of Medicine, 380, 2497-2505. DOI.
    15. Bodien, Y. G., Barra, A., Temkin, N. R., et al. (2024). Cognitive motor dissociation in disorders of consciousness. New England Journal of Medicine, 391, 598-608. DOI.
    16. Gallagher, S. (2008). Direct perception in the intersubjective context. Consciousness and Cognition, 17(2), 535-543. DOI.
    17. Butlin, P., Long, R., Elmoznino, E., et al. (2023). Consciousness in artificial intelligence: Insights from the science of consciousness. arXiv. DOI.
    18. Binder, F. J., Chua, J., Korbak, T., et al. (2025). Looking inward: Language models can learn about themselves by introspection. International Conference on Learning Representations. OpenReview.
    19. Song, S., Lederman, H., Hu, J., & Mahowald, K. (2025). Privileged self-access matters for introspection in AI. arXiv. DOI.
    20. Singh, S., Linzen, T., & Ravfogel, S. (2026). Can LLMs introspect? A reality check. arXiv. DOI.
    Evidence status and source ledger

    Established or replicated Testimony reliability is domain-sensitive; people track informant accuracy; self-explanations can diverge from manipulated causes; covert task responses can occur without observable command-following.

    Contested interpretation Whether ordinary understanding of others is inferential or direct; how much specific artificial-introspection results exceed learned cues and input-only prediction.

    Method proposed here The suspension bridge, source-dependence graph, counterfactual-selectivity rule and Testimony Warrant Record organise these results into one decision method.

    Open hypothesis Privileged internal access should transfer across surface changes and disappear under selective removal of the relevant access path.

    Category: Human and machine intelligence. Series: Inquiry and epistemic discipline. Taxonomy: Epistemology; philosophy of mind; consciousness science; AI evaluation; research methods. Tags: testimony, other minds, introspection, causal evidence, source dependence, cognitive motor dissociation, artificial consciousness.