Home · Writing · Consciousness

Emergence and the Ruler that Creates It

A measurement audit for deciding when a capability cliff is made by the metric, when it marks a dynamical transition and when an ontological claim has outrun the evidence.

TLDR

  1. A measurement audit for deciding when a capability cliff is made by the metric, when it marks a dynamical transition and when an ontological claim has outrun the evidence.
  2. At 9:00 on Monday morning, a laboratory tests ten related language models on eight-digit addition.
  3. The easiest way to misuse emergence is to leave its relata unspecified. What emerges, from what basis, for which observer, under which intervention and at what grain?
  4. M0 is a statement about an observation pipeline. M1 is a defensible product or laboratory statement when a capability really becomes usable.
  5. Three distinctions prevent a category error. Epistemic emergence concerns what an observer cannot predict or compress.
One capability coastline measured by four rulers A dark irregular coastline meets four differently spaced measurement grids. Fine, coarse, threshold and causal rulers produce different boundary summaries while the underlying coast remains unchanged. same underlying capability boundaryfine rulercoarse rulerthreshold rulercausal ruler pass line many small changesfewer larger stepsone apparent cliffintervention path
Figure 1. A ruler does not passively reveal a boundary. It selects resolution, threshold and relation. The coastline remains real, while its measured length and apparent discontinuities depend on how the observer samples it.
On this page

At 9:00 on Monday morning, a laboratory tests ten related language models on eight-digit addition. The first seven score zero under exact string match. The eighth gets 3 per cent, the ninth 27 per cent and the tenth 71 per cent. The chart looks like a wall. A new arithmetic ability seems to have appeared somewhere between models seven and eight.

At 9:20, the same outputs are scored again. Mean digit accuracy rises smoothly across all ten models. Token edit distance falls smoothly. The probability assigned to the correct answer rises smoothly. Even the small models were becoming less wrong, but the original ruler awarded them nothing until every digit landed in the right place.

At 10:00, a second surprise arrives. Across twenty training seeds, the middle-sized model separates into two groups. Half discover a compact addition routine; half remain with a brittle lookup strategy. The average score hides both populations. A larger model makes the compact routine much more likely. This transition is not removed by changing from exact match to edit distance. A hidden-state intervention disrupts the successful group in a way that it does not disrupt the lookup group.

Which event deserves the word emergence?

The first cliff belongs mainly to the ruler. The second may reflect a change in the distribution of learned mechanisms. Neither establishes strong emergence, and neither says that a new subject of experience came into being. The word has been asked to cover too many jobs: an observer’s threshold, a useful higher-level pattern, a dynamical phase change, a new mechanism and an ontologically novel power.

The central rule is simple: an emergence claim is only as strong as the transformations under which it remains true. Change the metric, sampling density, task mixture, random seed, prompt, scale coordinate and intervention. What survives those changes belongs increasingly to the system. What vanishes belongs increasingly to the measurement arrangement.

Part I. Five claims hiding inside one word

The easiest way to misuse emergence is to leave its relata unspecified. What emerges, from what basis, for which observer, under which intervention and at what grain? A flock shape can be unpredictable from casual inspection of one bird while remaining fully generated by local rules. Temperature is absent from one molecule yet indispensable at a thermodynamic scale. A benchmark skill can cross a product threshold without any abrupt change in the model. Consciousness raises a harder question because functional organisation and phenomenal presence may not share the same dependence relation.

The Stanford Encyclopedia account of emergent properties frames the broad idea through dependence on a lower-level basis together with some form of higher-level autonomy. That formulation is useful because novelty alone is cheap. Every arbitrary grouping creates a new description. The difficult question is whether the macro-description carries explanatory, predictive or causal work that cannot be recovered from the selected parts at the selected grain.

Emergence should be classified by the kind of autonomy claimed, not by how surprised the observer feels. Five classes keep the burdens separate without placing them on one ladder.

Class Claim Minimum evidence What it does not establish
M0, metric cliff A reported score changes sharply Exact scoring rule, outputs and uncertainty A new underlying ability
M1, operational onset A registered task becomes usable above a threshold Replication across samples, prompts and sensible metrics A new internal mechanism
M2, dynamical transition The system enters a distinct regime Order parameter, dense sweep, perturbation and preferably hysteresis Ontological novelty
M3, mechanistic reorganisation A different circuit or strategy causally supports performance Mechanistic identification and intervention Strong emergence or experience
M4, ontological emergence The whole has a power not exhausted by its physical basis A defended dependence relation, novel efficacy and exclusion analysis Automatic agreement across metaphysics

M0 is a statement about an observation pipeline. M1 is a defensible product or laboratory statement when a capability really becomes usable. M2 and M3 make different claims about the system’s internal organisation. M4 is a metaphysical claim. These classes can overlap, but none entails the next: an M3 mechanism may change smoothly without an M0 cliff or M2 phase transition, while an M2 transition need not cross an operational M1 threshold.

Weak emergence usually remains compatible with physicalism: macro-properties depend on and are realised by lower-level organisation, even when prediction requires simulation or a higher-level vocabulary. Strong emergence claims a deeper autonomy, often involving novel powers or laws not exhausted by the base. Calling a language-model score “strongly emergent” because a small model scored zero confuses five evidential burdens at once.

A higher-level description earns autonomy by compression, prediction or causal intervention, not by linguistic grandeur. A useful macro-variable may forecast the system better than any isolated component. A peer-reviewed study of causal emergence in multivariate systems formalises one version of that idea through collective information that influences future system states beyond the selected parts. Such a measure is a technical result relative to variables, times and a decomposition. It is not a certificate that the macro-property floats free of physics.

Conceptual depth: weak, strong and observer-relative emergence

Three distinctions prevent a category error. Epistemic emergence concerns what an observer cannot predict or compress. Weak ontological emergence grants real higher-level patterns while retaining determination by a physical basis. Strong emergence gives the higher level a more radical autonomy, sometimes including novel causal powers. These positions can agree on the same data and disagree about what the data mean.

Observer-relativity is not equivalent to unreality. A storm track depends on a chosen spatial and temporal grain, yet it can guide evacuation better than a molecular inventory. The test is whether the coarse-graining is stable, predictive and intervention-relevant for a declared purpose. Likewise, a capability boundary can be operationally real for deployment even when its apparent sharpness is produced by a pass threshold.

Part II. How a smooth ability becomes a cliff

Suppose a model’s probability of producing one correct digit is p(s), where s is a scale coordinate such as training loss, compute or effective data. The function can improve smoothly. Exact-match accuracy for an L-digit answer behaves approximately as p(s) raised to L when digit errors are independent. Raising a number smaller than one to the eighth or twentieth power hides early improvement and concentrates visible gains near the top of the range.

This is not a defect in exact match. A bank transfer reference, compiler output or arithmetic answer may need every symbol correct. Exact match answers an important operational question: did the whole artefact pass? It becomes misleading only when the operational score is treated as a direct meter of latent capability.

The same issue appears in multiple-choice grading. Imagine that the logit assigned to the correct answer rises smoothly from just below the leading distractor to just above it. Winner-take-all accuracy jumps from zero to one at the crossing. Brier score, log score and the correct-option margin register the approach before the winner changes. Again, both readings can be useful. They describe different properties.

A metric can be valid for release and invalid for mechanism discovery at the same time. The repair is not to discard hard pass conditions. It is to pair them with rulers that reveal how the system approaches, crosses and behaves beyond the boundary.

One rising tide read by three gauges A smooth teal tide rises across model scale. A probability gauge follows it smoothly, an exact-match gate stays closed then opens sharply, and a business threshold marks a real operational crossing without claiming a new mechanism. smallerlargerscale or falling training loss probability ruler follows the tide exact-match gate deployment threshold same model family, three legitimate questions How much probability moved?Did the whole answer pass?Is it usable here?
Figure 2. The probability ruler, exact-match gate and deployment threshold can all be legitimate. Confusion begins when a product boundary is redescribed as a sudden change in the model’s internal nature.

The most influential early catalogue defined large-model emergent abilities as abilities absent in smaller models and present in larger ones, with performance not predictable by extrapolating the smaller models. The TMLR survey of emergent abilities made the phenomenon visible across model families and tasks. It also created a precise target for criticism.

Schaeffer, Miranda and Koyejo later showed that many reported cliffs were attached to particular metrics rather than task-model pairs. In their NeurIPS metric analysis, more than 92 per cent of hand-annotated BIG-Bench emergence cases appeared under multiple-choice grade or exact string match. Continuous alternatives such as Brier score or token edit distance exposed graded improvement, and deliberately thresholded vision metrics could manufacture new-looking emergent abilities.

That result should neither be diluted nor universalised. It demonstrates a strong sufficient explanation for many cliffs. It does not prove that every capability changes smoothly under every useful coordinate. The paper itself leaves room for real emergence and analyses fixed outputs within model families. Training dynamics, random-seed mixtures and internal circuit changes remain open.

There are also two different ways for a smooth precursor to meet a discontinuous world. In the first, only the observer discretises: a probability of 0.49 and 0.51 becomes wrong then right under argmax. In the second, the environment itself contains a threshold. A compiler accepts or rejects a program, a tool either has a required permission, and a multi-step plan succeeds only if every dependency resolves. The system may improve smoothly while its consequences change sharply because the environment is conjunctive or irreversible.

That second case is still not evidence of a new internal mechanism, but it can create a real change in risk. If each of ten safety barriers fails with a small probability, correlated improvement or degradation can move the probability of joint failure nonlinearly. A deployment team should therefore keep the operational cliff while refusing to confuse it with a natural boundary in cognition. One curve governs product readiness; another supports scientific explanation.

Sparse observation can make either curve look more mysterious. Three checkpoints below a transition and one above it cannot distinguish a sigmoid, a kink, a discontinuity or a mixture of training outcomes. Logarithmic scale axes can visually compress large intervals, while aggregate means can place the apparent onset between points at which no individual seed changed. A claim of unpredictability needs an explicit forecasting exercise using only earlier checkpoints, not a retrospective impression from the completed curve.

Ruler What it rewards Why a cliff can appear Best companion measure
Exact string match Entire sequence correctness One error makes the score zero Token edit distance and sequence log probability
Multiple-choice accuracy Winning option Smooth logit crossing becomes a discrete flip Brier score, log score and option margin
Pass@1 First sampled solution succeeds Rare success is poorly resolved in small samples Pass@k, estimated success probability and uncertainty
Aggregate benchmark mean Average across heterogeneous items Offset, ceiling and mixture effects hide item curves Difficulty-conditioned response curves
Product acceptance gate Satisfies all constraints Conjunction multiplies failure probabilities Per-constraint risk and joint-failure model

Thought experiment: the ruler exchange

Imagine two sealed laboratories receive exactly the same model outputs. Laboratory A is told to score exact answers. Laboratory B is told to score normalised edit distance. A reports a capability onset at scale eight. B reports a smooth learning curve from scale two. Neither laboratory can inspect weights or run new generations.

Now exchange their rulers. Their conclusions exchange too. Nothing inside a model changed while the claims changed. This is a decisive diagnosis for M0 sharpness. It does not decide whether a later M2 or M3 transition exists, because neither laboratory has measured internal dynamics or intervened on a mechanism.

Smooth token improvement becomes an all-correct necklace cliff Four rows of eight beads represent answer tokens at increasing competence. More beads become correct smoothly, but a gold clasp labelled exact match closes only on the final all-correct row. eight-token answerexact-match score scale 1scale 2scale 3scale 4 0001 clasp closes only here correct tokenincorrect token
Figure 3. Per-token competence improves in the second and third rows, while exact match remains zero. The final score jump is operationally meaningful and mechanistically ambiguous.

The arithmetic is reproducible. The executable example below uses a smooth logistic token probability, then reads it through exact-match, expected edit accuracy and a winner-take-all multiple-choice gate.

from math import exp

def logistic(x: float) -> float:
    return 1.0 / (1.0 + exp(-x))

def rulers(scale: float, length: int = 8) -> dict[str, float]:
    token_p = logistic(0.75 * (scale - 5.0))
    correct_p = token_p
    distractor_p = 1.0 - correct_p
    return {
        "token_probability": token_p,
        "expected_token_accuracy": token_p,
        "exact_match_probability": pow(token_p, length),
        "multiple_choice_grade": float(correct_p > distractor_p),
        "two_class_brier": pow(correct_p - 1.0, 2) + pow(distractor_p, 2),
    }

rows = [rulers(scale) for scale in range(1, 10)]

assert all(rows[i]["token_probability"] < rows[i + 1]["token_probability"]
           for i in range(len(rows) - 1))
assert rows[4]["multiple_choice_grade"] == 0.0
assert rows[5]["multiple_choice_grade"] == 1.0
assert rows[2]["exact_match_probability"] < 0.001
assert rows[-1]["exact_match_probability"] > 0.65
assert all(rows[i]["two_class_brier"] > rows[i + 1]["two_class_brier"]
           for i in range(len(rows) - 1))
Mathematical depth: conjunctions sharpen without a phase transition

For a sequence of length L, exact-match probability is p^L under an independence approximation. Its derivative with respect to competence is L × p^(L-1). When p is modest, the derivative is tiny for long strings. Near one, it grows quickly. The metric therefore compresses early progress and expands late progress even when p itself follows a smooth curve.

Dependencies among token errors change the exact formula. They do not remove the general point. Any conjunctive score that requires many conditions to pass can create a steep response from smoother component probabilities. The correct analysis reports both the joint outcome and its component structure, with uncertainty and dependency estimates.

Part III. What survives a ruler change

Changing the metric is the first audit, not the last. A curve can acquire a false cliff through sparse scale points, small test sets, mixed item difficulty, prompt format, decoding policy, contamination, ceiling effects or averages across random seeds. Conversely, averaging can erase a real mixture transition. A smooth mean may combine two sharply different learned strategies.

The unit of analysis must therefore expand from model size to a measurement record:

R = (F, D, T, M, S, P, G, H)

Here F is the model family, D the training-data regime, T the task distribution, M the metric, S the sampling design, P the prompting and decoding policy, G the scale coordinate and H the intervention history. “Ability emerged at 10 billion parameters” suppresses nearly every term.

Model size is often a poor clock for capability. Training loss, effective compute, data quality and architecture can place different models at different functional stages despite similar parameter counts. A NeurIPS study of emergent abilities from the loss perspective reports that models with the same pretraining loss can show similar downstream performance under controlled corpus, tokenisation and architecture conditions, and also finds some task thresholds indexed by loss. That is evidence against a universal metric-mirage account and against parameter count as a sufficient coordinate.

Fine measurement can also expose progress below an apparent floor. The peer-reviewed ICLR PassUntil evaluation uses extensive sampling to give very small success probabilities measurable resolution. It finds task scaling that conventional pass rates miss and reports both predictable scaling and cases of accelerated improvement. The result does not restore every old cliff. It shows why “zero” can mean “below the test’s resolution”.

These results suggest a useful causal chain for an audit. Training changes a distribution over internal states and output probabilities. Decoding converts that distribution into sampled artefacts. A task parser decides what counts as a valid response. A metric maps valid responses into scores. Aggregation combines items, prompts, seeds and models. Plotting then selects axes, smoothing and scale. Every stage can introduce a threshold, and several thresholds can align.

The chain localises responsibility. If re-scoring fixed outputs removes the cliff, training and decoding are unchanged, so the sharpness entered at parsing, metric or aggregation. If the cliff survives re-scoring but moves under decoding temperature, the output distribution may be smooth while the sampling policy exposes it nonlinearly. If it survives ruler and decoding changes yet separates by training seed, the candidate explanation moves upstream toward learned strategy. If a targeted intervention abolishes the high-capability regime, the case for a mechanism becomes substantially stronger.

Forecasting should follow the same chain. Fit a continuous precursor on early checkpoints, propagate uncertainty through the registered operational metric and predict a distribution of possible crossing points. Then compare the observed crossing with that forecast. A capability can cross abruptly in product terms and still be forecastable from smooth precursors. Conversely, a continuous score may depart unexpectedly from its extrapolation without ever producing a visible step. Sharpness and unpredictability are independent properties and should be reported separately.

Rescoring and resolution extension are separate operations A registered archive containing outputs, probabilities and item identities enters a prism and supports four solid rescoring paths. A separate dashed PassUntil path returns to the generator for additional samples, making clear that it changes resolution rather than merely rescoring fixed outputs. registered archiveoutputs + probabilities + item IDs fixed archiverescored exact matchedit distanceBrieritem-response model generator + policy queried againPassUntil solid = rescore · dashed = add samples
Figure 4. Exact match, edit distance, Brier score and item-response analysis can use a sufficiently rich frozen archive. PassUntil is different: it extends decoding to resolve rare success. Both are useful, but only the solid paths isolate the effect of rescoring fixed evidence.

Worked example: four arithmetic releases

Consider four models tested on 1,000 eight-digit addition items with five prompts and ten random seeds. The figures below are illustrative audit fixtures, chosen to make the method inspectable rather than to report an external benchmark.

Release Mean digit accuracy Exact match Sequence log score Seeds using compact routine Targeted ablation δ exact match (matched random δ)
A 0.71 0.06 -4.82 0/10 -0.01 (-0.01)
B 0.79 0.15 -3.61 1/10 -0.03 (-0.02)
C 0.88 0.36 -2.31 6/10 -0.24 (-0.03)
D 0.94 0.61 -1.27 9/10 -0.29 (-0.03)

Exact match appears to accelerate from B to C. Digit accuracy and log score reveal earlier progress. Yet the seed-level strategy analysis adds information that metric replacement cannot explain: a compact routine becomes common, and its targeted ablation lowers exact match by 24 to 29 percentage points where matched random-state ablation lowers it by 3 points. The careful conclusion has two clauses. Part of the performance cliff is conjunctive scoring. A separate M3 mechanistic reorganisation is compatible with the controlled intervention and deserves replication across new seeds.

The design also prevents a common inference error. C is not a single deterministic phase point. It is a mixture of training outcomes. Saying “the model emerges at C” erases seed variance. The registered object is a distribution over trained systems under a specified recipe.

A scale by difficulty field with a threshold contour An illustrative smooth coloured field shows success probability rising with scale and falling with item difficulty. A dashed pass contour cuts the field and creates an apparent capability frontier. The geometry is qualitative and is not measured data. Illustrative · not measured harder itemsgreater effective scalechosen 70% pass contour item difficultysmalllarge low probabilityhigh
Figure 5. Illustrative qualitative schematic, not measured data. A pass contour can be operationally valuable without being a natural joint in capability space. A real study would estimate difficulty-conditioned curves and uncertainty rather than assume this geometry.
Survival test Hold fixed Change Evidence strengthened when
Metric exchange Outputs Exact, continuous, calibrated and decomposed scores Sharpness remains under several justified rulers
Resolution increase Task and model More items, samples and scale points The transition narrows rather than dissolves
Seed expansion Training recipe Random initialisation and data order A regime boundary appears in the outcome distribution
Coordinate exchange Models Parameters, compute, loss and effective data The boundary aligns with a mechanistically meaningful coordinate
Intervention test Baseline capability Ablate or perturb candidate mechanism The claimed regime depends selectively on the mechanism
Path reversal Endpoints Sweep control variable upward and downward Hysteresis or bistability appears reproducibly

No single survival test certifies emergence. Together they locate where the discontinuity lives. Metric exchange tests the ruler. Seed expansion tests the training distribution. Intervention tests the mechanism. Path reversal tests dynamical regime structure.

Evaluation depth: prompts, seeds and task mixtures

Prompting can shift a visible onset by supplying decomposition, demonstrations or output constraints. That is not necessarily cheating. It changes the system being evaluated from a bare model to a model-policy pair. The paper should name the pair and avoid comparing it with a different pair under one capability label.

Task mixtures can flatten or sharpen curves. Easy items may reach ceiling while difficult items remain at chance, causing the aggregate to stall. Later gains on difficult items then look sudden. Item-response models and difficulty bins expose this composition. Random seeds create another mixture: averages can place no actual trained model near the mean. Release claims should therefore include per-seed distributions and the probability of entering each regime.

Part IV. Brains, ignition and the ontological surcharge

Consciousness research supplies the hardest version of the ruler problem because the target is contested. A visual stimulus can vary continuously in contrast. Neural responses can be measured through spikes, local field potentials, scalp signals, imaging and report. Each channel has its own resolution and threshold. A participant can use a categorical response for an experience whose confidence or clarity is graded. A global broadcast can be abrupt while the local recurrent processing that feeds it is continuous.

Global neuronal workspace accounts use the idea of ignition: sufficiently strong information gains wide availability across a distributed network. Reviews of conscious processing and workspace dynamics motivate an all-or-none access transition, while empirical work also asks whether awareness is categorical or graded. A multisensory perceptual-awareness study found that the answer depends on what aspect and modality is measured. “Ignition” is therefore a mechanistic hypothesis with operational markers, not a synonym for the appearance of phenomenality.

A threshold in report can mark access, decision or motor commitment even when it is correlated with experience. No-report designs, graded confidence, trial-level neural dynamics and causal perturbation help separate these stages. They do not provide a view from outside every theory.

The ignition theatre separates latent evidence from observed thresholds A theatrical stage contains a smooth rising light field. Four curtains labelled local recurrence, global access, report and confidence open at different positions, producing different apparent ignition points. A separate upper balcony marks phenomenal presence as theory-mediated. smooth sensory evidence local recurrenceglobal accessreportconfidence phenomenal presence is not directly read from one curtain weaker stimulusstronger stimulusdifferent channels can create different ignition points
Figure 6. Neural recurrence, global availability, report and confidence can cross thresholds at different stimulus strengths. A consciousness theory must say which transition it treats as constitutive, evidential or incidental.

Phase-transition language can be more than metaphor when it comes with the mathematics and controls of dynamical systems. An order parameter changes across a control variable. Near a continuous transition, correlation length, susceptibility or variance may increase. A discontinuous transition can show jumps, bistability and hysteresis. Finite systems smear ideal singularities, so dense sampling and finite-size analysis matter.

Peer-reviewed models of stochastic spiking networks demonstrate both continuous and discontinuous transitions even when individual neurons have smooth firing probabilities. This is the important counterweight to the metric-mirage result: smooth units can collectively enter distinct regimes. Human brain recordings have also been analysed as occupying a continuum from second-order to first-order critical-like dynamics, including bistability-related signatures. These findings concern dynamical organisation. Their relationship to consciousness requires additional theory and contrastive evidence.

A hysteresis orbit distinguishes path-dependent regimes from a single score threshold An illustrative upward teal sweep and downward coral sweep follow different branches around a central bistable eye. The qualitative geometry demonstrates path dependence and is not measured data. Qualitative schematic · not data simple score threshold upward sweepdownward sweepcontrol variableorder parameterpath-dependent region
Figure 7. Illustrative qualitative schematic, not measured data. A score crossing a line is not a phase transition. A registered experiment would need to estimate forward and reverse paths, stable branches and uncertainty rather than assume the drawn loop.
Proposed signature What to measure Stronger interpretation requires Common confound
Sharp order-parameter change Registered macro-variable over dense control sweep Replication and finite-size analysis Sparse sampling or transformed axis
Critical slowing or rising variance Recovery time and fluctuation statistics Mechanistic link to the candidate transition Non-stationarity or measurement noise
Bistability Two stable regimes under matched controls State-dependent perturbation and dwell-time analysis Seed mixture treated as within-system dynamics
Hysteresis Different forward and reverse transition points Controlled path reversal Irreversible training history
Causal macro-variable Intervention or predictive value beyond selected parts Robustness across coarse-grainings Chosen variables encode the answer
Conscious access contrast Report, no-report and neural measures Theory-specific prediction and causal test Decision and motor thresholds

Criticality claims require particular restraint because power laws and avalanche-like distributions can arise from multiple mechanisms, finite samples and analysis choices. Evidence becomes more persuasive when several signatures converge under a generative model, when the proposed control parameter can be manipulated, and when alternative non-critical processes are compared directly. A straight line on log-log axes is an invitation to model comparison, not a phase-transition verdict.

The distinction between a phase change and a functionally important threshold also matters in brains. An organism can exploit a steep but continuous gain curve as if it were categorical. Recurrent amplification can create reliable access without an ideal thermodynamic singularity. Finite neural systems need not reproduce an infinite-system phase transition exactly for the mechanism to be useful. The scientific claim should match the finite system actually measured.

Likewise, a macro-variable can be causally informative without becoming an independent substance. If whole-network synchrony predicts recovery after perturbation better than any chosen neuron, that supports a higher-level description relative to those variables. It does not show that microphysics is incomplete. Causal-emergence measures are especially sensitive to the selected grain, the completeness of the component set and the time horizon. A robust study repeats the analysis across plausible partitions rather than celebrating the partition that maximises novelty.

For machine systems, the analogous test replaces neurons with components, activations, tokens or modules. A collective feature such as a persistent plan may predict later tool use beyond any one local state. The result can justify an agent-level control variable. It still leaves open whether the feature is implemented through distributed computation, external memory, orchestration or a measurement artefact. Emergence language should sharpen that investigation instead of ending it.

Does consciousness emerge?

A physicalist emergentist may say that experience depends on sufficiently organised matter and becomes a real higher-level property when that organisation is achieved. A strong emergentist may add novel powers or laws. A functionalist can locate the relevant transition in causal organisation. An illusionist can explain why a system represents itself as having ineffable properties without adding phenomenal properties to the ontology.

A consciousness-primary orientation reverses the verb. Awareness is not manufactured by complexity; organised systems delimit, express or localise a perspective within awareness. Śaṅkara’s non-dual tradition, surveyed in the Stanford Encyclopedia entry on Śaṅkara, treats consciousness as irreducible rather than as a late product. Process traditions in the West, represented broadly in process philosophy, replace inert substances with becoming and relation. These views do not generate a benchmark prediction merely by being named.

A consciousness-primary view removes the obligation to explain awareness from non-awareness, but it inherits a manifestation problem. Why does this organisation support a bounded perspective, memory continuity, report or suffering while another does not? Which changes alter the contents or boundary of a subject? A manifestation account must still expose conditions and contrasts. Otherwise it converts an explanatory gap into an unmeasured permission.

This is where the ruler audit remains useful across metaphysics. The physicalist asks whether an organisation produces experience. The consciousness-primary researcher asks whether it configures a locus of manifestation. Both should resist treating a benchmark cliff, an activation threshold or a verbal report as a complete answer.

Philosophical depth: emergence, manifestation and category mistakes

Aristotelian form and matter offer one historical way to describe a whole whose capacities are not a mere sum of detached parts. Modern weak emergence can preserve higher-level autonomy within physical determination. Strong emergence pays a larger price by challenging some version of causal closure. Non-dual and process-oriented traditions alter the background ontology more radically.

The traditions should not be collapsed. Saying that awareness is primary differs from saying that every organised system is a conscious subject. Saying that a property supervenes on a base differs from explaining why it exists. Saying that a macro-variable has unique predictive information differs from giving it fundamental causal power. Each move changes the burden rather than dissolving it.

Part V. The ruler-change protocol

An emergence result should ship with an audit record, not a dramatic curve alone. The protocol below is designed for model scaling, agent capability, neural access and other complex systems. It cannot force all domains into one statistic. It forces the claim to disclose which part of the result belongs to the system, the task and the observer.

First, register the claim class or classes. M0 asks whether a score jumps. M1 asks whether an ability becomes operationally usable. M2 asks whether a dynamical regime changes. M3 asks whether a new mechanism carries the capability. M4 asks for ontological autonomy. A study may support several classes, but evidence for one does not automatically count as evidence for another.

Second, freeze the raw record. Preserve item-level outputs, probabilities where available, sampling settings, prompts, seeds, checkpoints and failures. A chart is not raw evidence. Without the record, another evaluator cannot exchange rulers or estimate what the published aggregate concealed.

Third, build a ruler panel before looking for a cliff. Include the metric used for real-world acceptance, a continuous measure of proximity, a calibrated probabilistic score where the task permits it and a decomposition by item difficulty or subskill. Explain why every ruler is connected to the construct.

Fourth, increase resolution on both axes. Add scale points around the suspected transition and enough task samples to resolve rare success. Report uncertainty. A narrow confidence interval around zero is different from a zero caused by forty test items and one decoding attempt.

Fifth, separate population mixtures. Plot each training seed, architecture, prompt policy and meaningful task stratum. Use a hierarchical model when the design warrants one. Look for bimodality and strategy clusters rather than treating every mean curve as one representative system.

Sixth, identify an order parameter only if a dynamical claim is intended. The variable should distinguish regimes before it is fitted to the desired story. Test recovery, variance, state dwell time, forward and reverse paths, and sensitivity to finite system size.

Seventh, intervene on the proposed mechanism. Activation patching, ablation, state reset, pathway interruption, controlled noise or causal stimulation can distinguish a correlated signature from a carrier. Match the intervention to the grain of the claim.

Eighth, write the narrowest statement that survives. “Exact-match performance crosses our acceptance threshold” can be completely correct. “A compact arithmetic routine becomes more common across seeds and is causally necessary under this intervention” is stronger and different. “Reasoning emerged” may be too vague to evaluate.

Ninth, connect the result to a decision. A new operational onset may require access control, red-team coverage or revised forecasts even when it is metric-generated. A mechanistic transition may justify targeted monitoring or a change in training. A consciousness-related marker may justify precaution without licensing certainty about sentience. The decision and the ontology should be documented on separate lines.

Tenth, preserve negative outcomes. If a cliff dissolves under continuous scoring, publish the dissolution. If an apparent order parameter fails path reversal, retain that failure. If an ablation damages every task equally, it did not isolate the proposed mechanism. A research programme learns more from a well-localised non-result than from an emergence label that survives by changing its meaning.

For frontier-model governance, this protocol supports earlier warning. Continuous precursors can be monitored before a hard capability gate is crossed. Seed distributions can expose a rare dangerous strategy before the mean moves. Mechanistic probes can distinguish fluent imitation from a stable capability. None of these measures guarantees that the next training run will remain inside the observed family, but together they replace surprise as the default planning assumption.

The protocol does not ban the word emergence. It makes the word earn a stable referent.

An eight-stage astrolabe for ruler-change audits Eight coloured arcs on four guide rings form an astrolabe around a central claim. The arcs summarise claim class, raw record, ruler panel, resolution, mixtures, order parameter, intervention and release language. Decision linkage and preservation of negative results remain explicit steps in the prose. emergenceclaim 1 claim class2 raw record3 ruler panel4 resolution5 mixtures6 order parameter7 intervention8 release language alignment across rings strengthens the claim; one bright arc does not
Figure 8. The eight-stage astrolabe summarises the measurement core of the ten-step protocol. Decision linkage and preservation of negative results remain steps nine and ten in the prose; the visual does not replace them.

Language that preserves the evidence

Use M0 language when the result is metric-bound: “Exact-match accuracy rises sharply between the sampled checkpoints; token-level scores remain smooth.” Use M1 language for deployment: “The system crosses the registered acceptance boundary under this task distribution.” Use M2 language only after regime evidence: “The order parameter shows a replicated discontinuity with bistability under the registered sweep.” Use M3 language after intervention: “A newly prevalent circuit is necessary for the performance profile under this ablation.”

M4 language should state the metaphysical premise, the dependence relation and the causal claim. No benchmark can silently carry that surcharge. For consciousness, the conclusion must also separate evidence of report, access, integration, subject-boundary conditions and phenomenal presence.

The peer-reviewed TMLR BIG-Bench collaboration was valuable partly because it widened the task surface on which scale effects could be observed. Its descendants should widen the measurement surface too. The same principle applies to evaluations of agents. A binary “completed mission” score may hide smoothly improving planning, tool selection and recovery, while an operational system may still require the binary mission outcome. Preserve both.

Grokking offers another caution. In the original arXiv grokking preprint, generalisation can improve long after training performance saturates. That looks like delayed emergence, but the explanation depends on training dynamics, representation and regularisation rather than model size alone. It is a candidate M2 or M3 phenomenon only after the relevant state variables and mechanisms are tested. The name does not supply the phase theory.

Five independent claim classes orbit one study record A central study record connects independently to M0 metric cliff, M1 operational onset, M2 dynamical transition and M3 mechanism. No arrows connect these empirical classes. M4 ontology sits beyond a dashed bridge labelled explicit metaphysical premise. registeredstudy recordoutputs · dynamics · interventions M0metric cliffM1operational onsetM2dynamical transitionM3causal mechanism Explicit premise bridgeempirical record does not entail ontology M4ontology classes may co-occur · none is a prerequisite for another
Figure 9. M0 to M3 are independent empirical claim classes, not maturity levels. A mechanism can change smoothly, a dynamical transition can miss a product threshold, and an operational onset can be metric-made. Any move from the empirical record to M4 needs an explicit metaphysical premise.

Glossary

Term Working meaning in this paper
Ruler The metric, sampling design, resolution and transformation used to turn a system record into a reported result
Metric cliff A sharp change in a reported score that may arise from the measurement operator
Operational onset The point at which a capability satisfies a registered use condition
Order parameter A macro-variable chosen to distinguish dynamical regimes
Bifurcation A qualitative change in system dynamics as a control parameter varies
Hysteresis Path dependence in which forward and reverse sweeps change regime at different points
Mechanistic reorganisation A causally relevant change in the strategy, circuit or process supporting behaviour
Causal emergence Higher-order predictive or causal information not captured by selected components at a declared grain
Weak emergence Higher-level novelty or autonomy compatible with determination by a lower-level basis
Strong emergence A more radical claim of novel powers, laws or efficacy not exhausted by the physical basis

The decision this changes

Scaling decisions should no longer ask only, “Did a capability emerge?” They should ask four narrower questions. Did a registered score cross a threshold? Did the underlying output distribution change smoothly or discontinuously? Did the system enter a distinct dynamical or mechanistic regime? Which ontological conclusion, if any, is being added to those findings?

An invariance still separates observed curves from the claim that survives Five differently shaped observation curves enter three translucent curved filters labelled ruler exchange, resolution and intervention. Several lines bend or disappear. One teal invariant trajectory exits with its uncertainty band and a separate dotted premise path toward ontology. observed result familypublishable invariant ruler exchangeresolutionintervention claim + uncertainty ontology needs a separate premise lines that vanish still teach us where the apparent discontinuity entered
Figure 10. The final claim is the pattern that remains after justified transformations, with its uncertainty intact. Results that disappear are not wasted; they identify the ruler, resolution or mechanism on which the earlier story depended.

For production AI, this changes forecasting and control. Teams should monitor continuous precursor measures even when release depends on binary success. They should estimate the probability of rare capability across seeds and samples. They should not assume that a currently absent hard-pass score has no precursor, nor that every sharp score implies an unforecastable internal leap. Risk controls can be conservative without being conceptually careless.

For consciousness science, the protocol prevents three substitutions. Report is not silently substituted for experience. Neural ignition is not silently substituted for phenomenal onset. Organisational complexity is not silently substituted for ontological production. A physicalist, emergentist and consciousness-primary researcher can share the same audit record while stating different premise bridges.

The publishable result is the strongest invariant that remains after the ruler, resolution and intervention have changed. Sometimes that invariant will be a smooth capability curve with a real operational threshold. Sometimes it will be a reproducible phase transition or new mechanism. Sometimes the evidence will stop before ontology. That stopping point is a result, not a failure of imagination.

The word emergence becomes useful again when it no longer does the work of the missing experiment.