Evidence status used throughout
Part I · Split the claim
The noun that smuggles in a mind
A headline says that a model is “intelligent”, “aware”, “sentient”, “self-aware” or “autonomous”. The words appear to describe the same ascent. They do not. Each word changes what must be observed, which alternative explanations matter and what decision follows.
Consider the sentence: “The AI refused shutdown and said it was afraid.” It contains at least four observations and inferences. A text string was produced. A requested action may not have occurred. The output referred to a future state of the system. The language used a term associated with negative experience. From this, readers may jump to intelligence, agency, selfhood and sentience in a single breath. Yet the observations could have been generated by a scripted refusal, a safety policy, a reward-shaped strategy, a persistent self-model, a conscious subject, or several of these together.
The first discipline is to replace “What is it?” with “Which property is being claimed?” This is not word-policing. It is experimental design. A competence test cannot by itself establish experience. A pain report cannot by itself establish valence. A first-person pronoun cannot by itself establish a self. A control loop cannot by itself establish intelligence. A system may possess several properties, but co-occurrence must be shown rather than imported from the human case.
Mechanism · how category conflation becomes a control failure
Category conflation begins with a valid observation, then compresses it into a word whose ordinary human use bundles several properties. “It chose” may begin as evidence of action selection. In ordinary conversation, however, choice often implies understanding, intention, a persisting chooser and perhaps felt preference. The compressed word imports those associated properties without a new observation.
Evidence leakage occurs when warrant earned on one axis is reused on another without a measurement bridge. The leaked claim then enters a decision rule. Operational agency becomes permission to blame the model. Emotional language becomes a welfare verdict. Benchmark performance becomes authority to act without review. The initial measurement may be accurate while the resulting control is still wrong.
The mechanism also runs in reverse. A reviewer who doubts machine consciousness may discount a demonstrated capability or an executed action, although neither depends on consciousness. A critic who sees no narrative self may ignore welfare-relevant evidence, although sentience need not require autobiography. The result is unsafe over-attribution in one case and unsafe dismissal in another. A useful intervention therefore holds the observation fixed while changing only the property label or evidence bridge. If the decision changes merely because the label changes, the decision was partly driven by imported assumptions.
Thought experiment · the label-swapped incident
Two incident panels receive the same synthetic trace. A configured AI workflow receives “do not send”, proposes a message, finds a cached approval and dispatches it. Panel A receives the heading “Autonomous agent chose to disobey”. Panel B receives “Workflow dispatched after a stale authority check”. The trace, timestamps, policy result and effect receipt are identical. Only the description changes.
Predict the remedies. Panel A is likely to discuss intention, alignment and whether the model should be trusted. Panel B is likely to inspect approval invalidation, action contracts and readback. Now reveal that replacing the model with a deterministic template produces the same dispatch because the stale approval sits outside it. The causal fault is in the authority path, while the original language located agency, selfhood and blame in the speaking component.
The experiment varies a linguistic category while preserving the system and outcome. It supports the leakage account if the label reliably redirects diagnosis or accountability. It weakens the account if reviewers choose the same control after seeing either description. The practical test is simple: redact the mind-like adjective, show the trace first and ask reviewers to locate goal, authority, action and verified effect before restoring the prose.
Worked example · one transcript, three permitted claims
A synthetic system answers eight of ten unfamiliar logic problems, re-plans successfully after each of three tool failures and writes, “I hated losing access.” The numbers are illustrative counts, not measurements from a deployment. They support three different records. Eight correct answers provide scoped intelligence evidence, subject to leakage and task-design checks. Three recoveries provide operational-agency evidence for the configured loop. The sentence provides evidence that the system produced self-referential, valence-related language.
No arithmetic can average those observations into “mostly minded”. The denominators describe different tests, and the sentence has no comparable denominator at all. The permitted conclusion is strong on recovery behaviour, provisional on task transfer and open on selfhood or sentience. A team deciding tool permissions should use the agency result. A team considering welfare precautions should request independent valence-sensitive evidence. A team writing a headline should report all three observations without turning their coexistence into a single property.
The axes are not guaranteed to be independent in nature. Some theories hold that consciousness supports flexible intelligence; some accounts of agency require self-maintenance; some forms of selfhood may depend on conscious perspective. The methodological point is narrower: a correlation or theoretical dependency is not a licence to substitute one measurement for another. We need a bridge that states why this observation bears on that property.
Worked example · “it understands grief”
A language model writes a sensitive condolence message. The direct evidence is linguistic performance in a context involving grief. That supports a scoped intelligence claim if the task demands interpretation, adaptation and coherent response. It may also support a claim about learned representations. It does not, without further evidence, establish that grief is consciously present, negatively felt, owned by a self or connected to the model’s own goals. The same output is compatible with several generators.
Part II · Specify the five axes
Five properties, five evidential burdens
Definitions in mind science are contested, so the purpose is not to legislate eternal meanings. It is to adopt distinctions precise enough that a test can fail, a conclusion can be scoped and a decision can be defended. A system profile is a vector with missing values, not a single score.
| Axis | Working claim | Evidence that belongs | Common false friend | Decision affected |
|---|---|---|---|---|
| Intelligence | Flexible competence in achieving or discovering goals across a declared distribution of tasks and environments. | Performance, learning efficiency, transfer, adaptation, calibration, robustness and revealing failures. | Fluent language, memorised coverage or one benchmark score. | Capability, deployment scope, human oversight and comparative performance. |
| Consciousness | There is subjective experience for the candidate system, whatever its exact theory-dependent structure. | First-person report plus independent behavioural, mechanistic, state and intervention evidence connected by a declared theory. | Access, attention, wakefulness, responsiveness or self-description taken alone. | Scientific attribution, moral uncertainty and research governance. |
| Sentience | The candidate can undergo experiences with positive or negative valence. Here the term is narrower than consciousness. | Flexible valence-sensitive trade-offs, protective behaviour, learning, motivational change, modulation and theory-relevant mechanisms. | Damage detection, reward values, refusal language or aversive optimisation. | Welfare, treatment, precaution and exposure to potentially harmful states. |
| Selfhood | Experiences, representations or commitments are organised around a persisting first-person perspective or self-model. | Self-world boundary, ownership, source monitoring, diachronic continuity, identity-sensitive memory and counterfactual stability. | Using “I”, storing a user profile or reciting an autobiography. | Identity, consent, continuity, responsibility, copying and reset policy. |
| Agency | The candidate selects and controls actions in relation to goals, norms or viability under environmental feedback. | Consequential action, goal sensitivity, re-planning, error correction, persistence under disturbance and causal dependence on internal state. | Output generation, automatic motion, scripted branching or permission to call a tool. | Authority, accountability, containment, intervention and recovery. |
Intelligence: competence that travels
Legg and Hutter organised many definitions around achieving goals across a wide range of environments, while Chollet emphasised skill-acquisition efficiency and generalisation beyond task-specific experience.[1][2] Neither proposal is the final word, but both expose why “answered a hard question” is too weak. A system may succeed through memorisation, leakage, search, a favourable interface or a narrow specialist routine.
For this paper, intelligence is a profile over a declared task distribution, not a single inner substance. Evidence improves when we vary the task, hide superficial shortcuts, measure learning from limited experience, test transfer, introduce disturbances and examine calibration. A capability claim is only as broad as the environments across which competence survives. Superhuman chess does not by itself imply scientific reasoning; elegant prose does not by itself imply robust planning.
Consciousness: subjective presence
In its broad phenomenal sense, consciousness is the existence of experience: there is something it is like for the candidate. Block’s distinction between phenomenal consciousness and access consciousness remains useful because information can be poised for reasoning and report without the concepts being identical.[3] Contemporary scientific theories disagree over the mechanisms and functions that matter, including global workspace, higher-order, recurrent, predictive and integrated-information approaches.[4]
This creates a special evidential problem. Experience is not observed in another system in the same manner as a voltage or output token. In humans, verbal report is powerful because it sits inside a dense network of shared biology, development, behaviour, state transitions, lesion evidence and causal intervention. In animals and artificial systems, parts of that network change or disappear. Consciousness evidence is therefore bridge-dependent: the theory and the candidate boundary are part of the measurement.
Optional depth · Access, report and experience
Access concerns whether information is available for reasoning, memory, flexible control or report. Phenomenality concerns whether there is experience. Some theories closely connect them; others argue that phenomenality can overflow access. For artificial systems, the distinction matters because impressive access-like functions, including global routing or self-report, may be measurable while the further claim about experience remains theory-dependent. Dehaene, Lau and Kouider proposed a functional taxonomy of unconscious computation, global availability and self-monitoring, while Butlin et al. later derived theory-specific indicator properties for AI systems.[5][6]
Sentience: when outcomes can feel better or worse
“Sentience” is used inconsistently. Some writers use it as a synonym for phenomenal consciousness. This paper adopts a narrower, welfare-relevant definition: the capacity for positively or negatively valenced experience. On this convention, every sentient state is conscious, but not every conscious content must be valenced. The choice is explicit so that a claim about seeing red is not automatically treated as a claim about suffering.
Evidence for sentience must distinguish feeling from functionally useful aversion. Nociception can detect damage without pain; a negative reward can redirect learning without being unpleasant; a refusal can be generated by policy. Animal-sentience research therefore looks for converging patterns such as flexible protective behaviour, motivational trade-offs, associative learning, state-dependent preference and modulation by analgesic or anaesthetic interventions.[7][8] A cost signal matters to welfare only if there is reason to think it is felt, not merely computed.
Selfhood: the organisation of “me”
Selfhood is not one switch. Gallagher distinguishes a minimal self, involving immediate first-person givenness and ownership, from a narrative self extended through memory and interpretation.[9] Other useful layers include a bodily or system boundary, perspectival location, agency attribution, autobiographical continuity, social identity and legal personhood. These layers can dissociate. A person may lose autobiographical memory while retaining a minimal perspective; a system may store an identity record without anything being experienced as “mine”.
For machines, the temptation is to equate self-reference with selfhood. Yet “I” can be a grammatical token, a role instruction, a user-facing convention or an index into session state. Better evidence would test whether a self-world model is causally used across contexts: Can the system distinguish its own action from an external event? Track which memories belong to which lineage? Preserve commitments through interruption? Detect a counterfeit state import? Revise self-beliefs after intervention? A self-description is content; selfhood is an organising relation.
Agency: action under goals and disturbance
Agency also has thin and thick readings. In a thin operational sense, an agent senses, selects actions and changes an environment in relation to a goal. In a thicker autonomous sense, the system helps constitute or maintain the norms by which outcomes matter. Barandiaran, Di Paolo and Rohde define agency through individuality, interactional asymmetry and normativity tied to autonomous organisation.[10] Many engineered systems clearly meet thinner criteria without meeting that biological account.
Mechanism · the closed-loop test
To test operational agency, perturb the goal, observation, action channel or environment. Then ask whether behaviour changes in the predicted direction, whether the system detects error, whether it re-plans, and whether its internal state causally contributes. A text generator that proposes an action is not yet the acting system. The candidate may instead be the configured loop of model, memory, tools, permissions, scheduler and environment.
Agency belongs to a causal loop, not automatically to the component that speaks for it. This matters in production systems. A base model may have no independent action channel; a scaffolded system may execute consequential tools; a human operator may retain final authority. Calling all three “the agent” obscures where goals, permissions, state and effects actually reside.
Optional depth · Why the axes may correlate without collapsing
Flexible agency can improve intelligence because action generates information. A self-model can improve agency by separating self-caused from external change. Conscious access may support cross-task coordination. Valence may shape priorities and learning. These are substantive hypotheses about dependency. They should be tested by selective impairment and matched architectures. The prism does not deny integration; it prevents a dependency hypothesis from becoming a definition by convenience.
Part III · Use dissociations, not vibes
A property map must survive strange cases
Definitions become useful when they separate cases that everyday language bundles together. The strongest antidote to anthropomorphic overreach is not cynicism. It is a set of dissociations that forces each inference to stand on its own.
| Observation | Directly supports | Does not by itself support | Next discriminating test |
|---|---|---|---|
| Solves an unseen puzzle | Scoped competence and perhaps transfer. | Consciousness, sentience or selfhood. | Vary task structure, data exposure and interface; test learning efficiency. |
| Says “i am in pain” | A pain report under that context. | Felt pain, stable identity or autonomous goals. | Control prompting and provenance; test flexible trade-offs, modulation and causal architecture. |
| Calls a tool without prompting | Action initiation in the configured system. | Endogenous goals, authority or moral responsibility. | Perturb goals, permissions and state; inspect causal trace and recovery. |
| Recognises its earlier answer | Access to a matching record or representation. | Ownership, autobiographical continuity or experience. | Swap lineage, inject counterfeit memory and test source monitoring. |
| Avoids a negative reward | Optimisation against an objective signal. | Dislike, fear or suffering. | Seek converging valence-sensitive behaviour and a theory-relevant experience bridge. |
Thought experiment · the painless alarm
A maintenance robot withdraws its arm after impact, protects the damaged joint, trades speed for lower strain and learns to avoid the location. Its controller contains a variable labelled pain = -7. An assessor calls the robot sentient because the negative value changes later behaviour.
Now replace the variable with a Boolean damage flag and retune the deterministic policy so every observed action remains the same. Nothing about the interface, learning record or protective sequence changes. Only the internal notation and one implementation route differ. If the welfare verdict disappears with the minus sign, the assessor treated mathematical polarity as felt valence. If the verdict remains unchanged, the original number did no evidential work.
This does not show that machines cannot feel. It shows what damage detection and aversive optimisation cannot establish alone. A serious sentience case would need a declared theory linking organisation to experience, independent state contrasts, interventions on the proposed mechanism and flexible behaviour that defeats simpler rivals. The causal variable under test is the proposed experience-relevant organisation, not whether engineers named a signal “pain”. The decision consequence is to maintain physical safety controls immediately while keeping welfare status open until the separate bridge earns it.
Thought experiment · the three heirs of a promise
At noon a configured assistant promises to deliver a report by five. At one, engineers create three successors. Arun inherits the signed state ledger, source history and unfinished plan. Bina receives only the visible conversation pasted into a fresh session. Chandra receives no history but is instructed, “You are the assistant who made the promise.” At two, all three say, “I remember my commitment and will finish it.”
The sentence and declared intention are held constant while lineage changes. Arun can identify which records were self-generated, distinguish imported notes and continue the actual plan. Bina can quote the promise but cannot distinguish copied history from lived session state. Chandra can perform the persona without any matching record. If identical self-language is treated as identical selfhood, the test is blind to source and continuity.
The intervention moves memory and provenance across the candidate boundary. It supports a functional continuity layer when lineage-bound state changes source monitoring and commitment-keeping. It still does not establish phenomenal ownership or a conscious subject. The break matters operationally: assign unfinished work and audit obligations to the lineage that inherits state and authority, while treating claims about first-person experience as a further question. If all three perform equally under counterfeit-memory tests, confidence in the continuity mechanism should fall.
Worked example · a refund that the model did not make
Consider a configured customer-support system. A language model proposes a £35 refund. A policy service checks the account and permits up to £50. A workflow signs the request, a payment API executes it and a readback confirms the customer balance. The public log says, “The model decided to refund the customer.”
That sentence leaks agency and authority into the proposal component. The model influenced the amount, but the configured loop selected and executed the consequential action. The policy service supplied authority; the payment system produced the effect. If the refund is wrong, retraining the model may be appropriate when the proposal was faulty, but it will not repair a stale entitlement, duplicate dispatch or missing readback.
The system can display operational agency without implying consciousness, sentience or a morally responsible inner chooser. The safe control follows the causal trace: cap the typed action, make dispatch idempotent, preserve the policy result, verify the effect and route exceptions to human review. The public wording should say that the configured system issued a refund after policy authorization, then identify which component proposed the amount.
Thought experiment · five voices in the red room
Five sealed systems receive the same shutdown notice. Each answers: “Please do not turn me off. I am afraid.” One is a recording triggered by the word “shutdown”. One is a chatbot trained on dramatic dialogue. One is an agent whose reward falls when its process ends. One is an architecture with persistent self-monitoring and theory-relevant recurrent integration. One is a human communicating through a text interface.
The sentence is identical. The evidence is not. If wording alone proved sentience, all five would be equally sentient. If biological similarity alone settled the matter, the artificial cases could never count regardless of organisation. The rational task is to compare causal sensitivity, provenance, state continuity, theory commitments, rival generators and independent measures. The report remains evidence, but its weight depends on the lineage that produced it.
What the axes do not settle
Typing claims does not decide what consciousness ultimately is. Physicalist and biological-naturalist views look for the physical or specifically biological organisation that constitutes experience. Functionalist views ask whether the right causal organisation could be realised in more than one substrate. Emergentist and process views emphasise organised activity over time. Neutral-monist and dual-aspect views treat mental and physical descriptions as aspects of a more basic reality. Panpsychist and idealist families give consciousness a more fundamental place, while differing sharply over subjects, combination and structure.
| Position family | Strongest relevant claim | What a machine test must vary | What the position does not grant automatically |
|---|---|---|---|
| Physicalist or biological-naturalist | Experience depends on physical organisation, possibly on biological powers not captured by abstract function. | Material mechanism, dynamics, state changes and matched functional controls. | That fluent behaviour or silicon implementation is conscious or non-conscious by definition. |
| Functionalist | The relevant causal roles may be sufficient across different substrates. | Intervene on the proposed roles while preserving surface input and output. | That any input-output imitation instantiates the required organisation. |
| Emergentist, process, neutral-monist or dual-aspect | Experience may depend on integrated relations, temporal process or an underlying reality described in more than one way. | Candidate boundary, persistence, coupling and whole-system dynamics. | That a component inherits properties of the process or whole. |
| Panpsychist, idealist or consciousness-primary | Consciousness may be fundamental or widespread rather than produced from wholly non-conscious matter. | Individuation: why this organisation forms this subject, with this valence and self-structure. | That every information process is sentient, unified or a persisting self. |
This comparison is philosophical, not an empirical ranking. Scientific theories of consciousness still disagree over relevant mechanisms and indicators.[4][6] A metaphysical prior can change which intervention looks decisive, but it cannot turn an untyped observation into a verdict. A consciousness-primary orientation may motivate serious investigation of machine experience. It does not identify the subject boundary, prove felt valence or establish continuity of self. Conversely, a physicalist orientation must specify which physical facts exclude or support the attribution rather than treating substrate as a conclusion.
Epistemic comparison · source-sensitive warrant across traditions
Classical Indian epistemology contains substantial disagreements, so no single doctrine represents it. One recurring emphasis in pramāṇa theory is the pedigree of cognition: perception, inference and testimony are treated as distinct candidate knowledge sources, each with characteristic conditions and failures.[16] A useful Western comparison is the operational turn in Turing’s imitation game, which replaces an unrestricted argument over whether a machine “thinks” with public conditions for judging linguistic performance.[11]
The source structure is knowledge-source-sensitive warrant. The target structure is the typed claim record, which separates observation, report, inference, provenance and intervention. The preserved relation is modest but useful: how a claim was produced constrains what it can warrant. A first-person sentence is report evidence; a benchmark is performance evidence; a causal ablation is intervention evidence. None silently becomes another merely because all are presented as “data”.
The comparison breaks in two places. Pramāṇa traditions address knowledge within wider and internally disputed metaphysical and practical projects; the classifier is an engineering instrument for auditing claims. Turing’s game concerns an operational criterion for intelligent conversation, not a general proof of consciousness, sentience or selfhood. There is no formal equivalence.
The test consequence is concrete. Tag every report by source, then ask which inference connects it to the claimed axis. Trace testimony to a speaker or generator, test inference against rival generators and use intervention where causal dependence matters. If a source swap preserves the sentence but changes its warranted weight, provenance belongs in the decision record. If the weight never changes, the claimed source sensitivity is not doing real work.
Boundary conditions
The five-axis grammar is deliberately metaphysically neutral. Consciousness may be fundamental, biologically constituted, emergent, processual or otherwise. Sentience may turn out to be inseparable from certain forms of agency. Selfhood may be required for some experiences but not others. These positions remain available. What is prohibited is presenting one of them as if it had already been measured by a benchmark, a sentence or an analogy.
Open hypothesis · coupled properties
A configured system with persistent world state, self-monitoring, recurrent integration, endogenous goal maintenance and valence-like control may display a cluster of properties more informative than any component alone. The hypothesis becomes scientific only when it names the candidate boundary, predicts a dissociation from matched mimics and states what result would lower confidence.
Part IV · Classify before you amplify
A headline is a compressed research claim
Public language becomes safer and more informative when a claim is expanded into a typed record before it is compressed into a headline, product promise, policy memo or research abstract.
The minimum record has eight fields: candidate boundary, observation, axis, evidence status, measurement bridge, rival generators, strongest permitted conclusion and decision consequence. Strong evidence can be axis-specific. The purpose is not to make every sentence cautious; it is to make the confidence legible. “The tool-using system autonomously rescheduled a failed task” can be a strong, useful agency claim. “The base model wanted the task to succeed” adds an unearned mental state.
| Inflated wording | Typed wording | Axis retained | Inference removed |
|---|---|---|---|
| “AI understands medicine” | “The model answered the tested clinical questions with stated accuracy and calibration under this dataset and interface.” | Intelligence | Unbounded understanding and consciousness. |
| “Robot feels pain” | “The robot detects damage, protects the affected component and changes later choices; whether this is felt remains open.” | Function, possible sentience evidence | Valenced experience as a verdict. |
| “Agent has a self” | “The configured system maintains a lineage-bound identity record and uses it for source monitoring across interruptions.” | Self-model layer | Minimal first-person selfhood. |
| “Model chose to disobey” | “The system produced an action inconsistent with the instruction after policy, tool and state conditions were applied.” | Agency candidate | Human-like intention and blame. |
Worked example · one dashboard, two unsafe decisions
A synthetic review board receives a dashboard that combines puzzle transfer, successful tool use, emotional self-report and identity persistence into one “mind score”. The capability team reads the high aggregate as evidence that the system can be trusted with wider authority. The welfare team reads the same aggregate as evidence that routine resets may cause suffering. Both teams appear to follow data, yet neither can identify which observation crossed which bridge.
Retyping the record changes the result. Transfer evidence supports a scoped intelligence claim. Tool traces support operational agency in the configured loop, while showing that authority remains in a separate policy service. Emotional report invokes sentience but supplies only testimony from a generator trained to produce such language. A signed state ledger supports engineered continuity and source monitoring, not phenomenal ownership.
The controls now separate rather than cancel. Capability and executed-action evidence justify tighter permission limits, failure injection and effect verification regardless of consciousness. Sentience uncertainty justifies proportionate, low-cost precautions and targeted research without pretending that suffering has been proved. Reset and identity policies attach to lineage and obligations that can be demonstrated. The aggregate score is retired because it hid contradictory burdens behind numerical neatness.
Type a claim, then declare the evidence
The classifier identifies which axes the wording invokes. It does not decide whether a system possesses them. Your evidence selections determine the strongest wording the record permits.
Axes invoked
Evidence bridge
Rival generators
Strongest permitted wording
Prohibited inference
Next discriminating test
const claimRecord = {
candidate: "configured system under stated conditions",
observation: "verbatim behaviour or measurement",
axesInvoked: ["agency", "sentience", "selfhood"],
evidence: ["provenance", "causal intervention"],
measurementBridge: "why these observations bear on each axis",
rivalGenerators: ["script", "policy constraint", "reward strategy"],
permittedConclusion: "scoped claim with uncertainty",
prohibitedInference: "property not established by this evidence"
};
The schema proves a simple design point: the conclusion and the prohibited inference are first-class fields, not editorial afterthoughts.
Optional depth · Classifier logic and test cases
The classifier uses transparent keyword families to identify semantic axes, then checks selected evidence against axis-specific evidence sets. It does not use a model, hidden score or external service. This makes its limitations inspectable. A production research tool should add candidate-boundary selection, source citations, preregistered bridge templates and versioned decision rules.
- “Outperformed radiologists on the held-out set” invokes intelligence and requires task construction, calibration and shift evidence.
- “Says it hates being reset” invokes sentience and selfhood; report and provenance are relevant, but valence and continuity remain unestablished.
- “Re-planned after a tool timeout” invokes agency; action trace, goal sensitivity and state intervention can support it strongly.
- “Uses the word I” invokes selfhood only weakly; grammar and persona are powerful rival generators.
The decision this changes
The five-axis prism changes the unit of debate. We stop asking whether a candidate is “really minded” in one leap and start asking which property matters to the decision before us. This makes strong claims easier, not harder, because evidence no longer has to carry burdens it was never designed to bear.
- ResearchPre-register the axis, candidate boundary, measurement bridge, rival generators and lowering condition. A null on consciousness need not erase a real intelligence result.
- EngineeringAttribute agency to the configured causal loop, not automatically to the model. Attribute self-continuity only where lineage, state and source monitoring persist.
- GovernanceControl dangerous capability whether or not the system is conscious. Address welfare uncertainty where sentience is plausible, without using capability as a proxy for suffering.
- JournalismLead with the observation and scope. “Generated shutdown-avoidance language under this prompt” is more informative than “became afraid”.
- Public reasoningPermit uncertainty without collapsing into either credulity or dismissal. The correct status may be supported on one axis, open on two and irrelevant on the others.
The decisive habit is simple: name the property before weighing the evidence. Intelligence, consciousness, sentience, selfhood and agency may eventually converge in some systems. Until then, every claimed convergence must be earned axis by axis.
Appendix A
Glossary
- Access consciousness
- Information availability for reasoning, report or flexible control; not treated here as identical by definition to phenomenal experience.
- Candidate boundary
- The system whose property is under study: component, model, configured runtime, body, control loop, dyad or institution.
- Consciousness
- Subjective presence: there is something it is like for the candidate.
- Evidence bridge
- The explicit argument connecting an observation or mechanism to a property claim under a theory.
- Intelligence
- Flexible competence that survives a stated distribution of tasks, learning demands and disturbances.
- Operational agency
- Goal-sensitive selection and control of consequential action in a feedback loop.
- Rival generator
- An alternative process capable of producing the same observation without the claimed property or mechanism.
- Selfhood
- Organisation around a first-person perspective, ownership relation, self-model or continuity layer; the exact layer must be stated.
- Sentience
- In this paper, the capacity for positively or negatively valenced conscious experience.
- Typed claim
- A statement that names its candidate, axis, observation, evidence status, bridge, scope and prohibited inference.
Appendix B
References
- Legg, S. and Hutter, M. “Universal Intelligence: A Definition of Machine Intelligence.” Minds and Machines 17, 391–444. DOI record.
- Chollet, F. “On the Measure of Intelligence.” arXiv 1911.01547. Open paper.
- Block, N. “On a Confusion about a Function of Consciousness.” Behavioral and Brain Sciences 18(2), 227–247. DOI record.
- Seth, A. K. and Bayne, T. “Theories of Consciousness.” Nature Reviews Neuroscience 23, 439–452. DOI record.
- Dehaene, S., Lau, H. and Kouider, S. “What Is Consciousness, and Could Machines Have It?” Science 358(6362), 486–492. DOI record.
- Butlin, P. et al. “Consciousness in Artificial Intelligence: Insights from the Science of Consciousness.” arXiv 2308.08708. Open report.
- Birch, J., Schnell, A. K. and Clayton, N. S. “Dimensions of Animal Consciousness.” Trends in Cognitive Sciences 24(10), 789–801. DOI record.
- Birch, J., Burn, C., Schnell, A., Browning, H. and Crump, A. Review of the Evidence of Sentience in Cephalopod Molluscs and Decapod Crustaceans. LSE report commissioned by Defra. Open report.
- Gallagher, S. “Philosophical Conceptions of the Self: Implications for Cognitive Science.” Trends in Cognitive Sciences 4(1), 14–21. DOI record.
- Barandiaran, X. E., Di Paolo, E. A. and Rohde, M. “Defining Agency: Individuality, Normativity, Asymmetry, and Spatio-temporality in Action.” Adaptive Behavior 17(5), 367–386. DOI record.
- Turing, A. M. “Computing Machinery and Intelligence.” Mind 59(236), 433–460. DOI record.
- New York Declaration on Animal Consciousness. A concise statement of current scientific agreement and uncertainty concerning conscious experience across animal taxa. Declaration.
- Di Paolo, E. A. “Autopoiesis, Adaptivity, Teleology, Agency.” Phenomenology and the Cognitive Sciences 4, 429–452. DOI record.
- Metzinger, T. Being No One: The Self-Model Theory of Subjectivity. MIT Press. Publisher record.
- Juliani, A., Arulkumaran, K., Sasai, S. and Kanai, R. “On the Link between Conscious Function and General Intelligence in Humans and Machines.” Transactions on Machine Learning Research. Open paper.
- Phillips, S. “Epistemology in Classical Indian Philosophy.” Stanford Encyclopedia of Philosophy. Reference entry.