At 9:00 on Monday morning, a laboratory tests ten related language models on eight-digit addition. The first seven score zero under exact string match. The eighth gets 3 per cent, the ninth 27 per cent and the tenth 71 per cent. The chart looks like a wall. A new arithmetic ability seems to have appeared somewhere between models seven and eight.
At 9:20, the same outputs are scored again. Mean digit accuracy rises smoothly across all ten models. Token edit distance falls smoothly. The probability assigned to the correct answer rises smoothly. Even the small models were becoming less wrong, but the original ruler awarded them nothing until every digit landed in the right place.
At 10:00, a second surprise arrives. Across twenty training seeds, the middle-sized model separates into two groups. Half discover a compact addition routine; half remain with a brittle lookup strategy. The average score hides both populations. A larger model makes the compact routine much more likely. This transition is not removed by changing from exact match to edit distance. A hidden-state intervention disrupts the successful group in a way that it does not disrupt the lookup group.
Which event deserves the word emergence?
The first cliff belongs mainly to the ruler. The second may reflect a change in the distribution of learned mechanisms. Neither establishes strong emergence, and neither says that a new subject of experience came into being. The word has been asked to cover too many jobs: an observer’s threshold, a useful higher-level pattern, a dynamical phase change, a new mechanism and an ontologically novel power.
The central rule is simple: an emergence claim is only as strong as the transformations under which it remains true. Change the metric, sampling density, task mixture, random seed, prompt, scale coordinate and intervention. What survives those changes belongs increasingly to the system. What vanishes belongs increasingly to the measurement arrangement.
Part I. Five claims hiding inside one word
The easiest way to misuse emergence is to leave its relata unspecified. What emerges, from what basis, for which observer, under which intervention and at what grain? A flock shape can be unpredictable from casual inspection of one bird while remaining fully generated by local rules. Temperature is absent from one molecule yet indispensable at a thermodynamic scale. A benchmark skill can cross a product threshold without any abrupt change in the model. Consciousness raises a harder question because functional organisation and phenomenal presence may not share the same dependence relation.
The Stanford Encyclopedia account of emergent properties frames the broad idea through dependence on a lower-level basis together with some form of higher-level autonomy. That formulation is useful because novelty alone is cheap. Every arbitrary grouping creates a new description. The difficult question is whether the macro-description carries explanatory, predictive or causal work that cannot be recovered from the selected parts at the selected grain.
Emergence should be classified by the kind of autonomy claimed, not by how surprised the observer feels. Five classes keep the burdens separate without placing them on one ladder.
| Class | Claim | Minimum evidence | What it does not establish |
|---|---|---|---|
| M0, metric cliff | A reported score changes sharply | Exact scoring rule, outputs and uncertainty | A new underlying ability |
| M1, operational onset | A registered task becomes usable above a threshold | Replication across samples, prompts and sensible metrics | A new internal mechanism |
| M2, dynamical transition | The system enters a distinct regime | Order parameter, dense sweep, perturbation and preferably hysteresis | Ontological novelty |
| M3, mechanistic reorganisation | A different circuit or strategy causally supports performance | Mechanistic identification and intervention | Strong emergence or experience |
| M4, ontological emergence | The whole has a power not exhausted by its physical basis | A defended dependence relation, novel efficacy and exclusion analysis | Automatic agreement across metaphysics |
M0 is a statement about an observation pipeline. M1 is a defensible product or laboratory statement when a capability really becomes usable. M2 and M3 make different claims about the system’s internal organisation. M4 is a metaphysical claim. These classes can overlap, but none entails the next: an M3 mechanism may change smoothly without an M0 cliff or M2 phase transition, while an M2 transition need not cross an operational M1 threshold.
Weak emergence usually remains compatible with physicalism: macro-properties depend on and are realised by lower-level organisation, even when prediction requires simulation or a higher-level vocabulary. Strong emergence claims a deeper autonomy, often involving novel powers or laws not exhausted by the base. Calling a language-model score “strongly emergent” because a small model scored zero confuses five evidential burdens at once.
A higher-level description earns autonomy by compression, prediction or causal intervention, not by linguistic grandeur. A useful macro-variable may forecast the system better than any isolated component. A peer-reviewed study of causal emergence in multivariate systems formalises one version of that idea through collective information that influences future system states beyond the selected parts. Such a measure is a technical result relative to variables, times and a decomposition. It is not a certificate that the macro-property floats free of physics.
Conceptual depth: weak, strong and observer-relative emergence
Three distinctions prevent a category error. Epistemic emergence concerns what an observer cannot predict or compress. Weak ontological emergence grants real higher-level patterns while retaining determination by a physical basis. Strong emergence gives the higher level a more radical autonomy, sometimes including novel causal powers. These positions can agree on the same data and disagree about what the data mean.
Observer-relativity is not equivalent to unreality. A storm track depends on a chosen spatial and temporal grain, yet it can guide evacuation better than a molecular inventory. The test is whether the coarse-graining is stable, predictive and intervention-relevant for a declared purpose. Likewise, a capability boundary can be operationally real for deployment even when its apparent sharpness is produced by a pass threshold.
Part II. How a smooth ability becomes a cliff
Suppose a model’s probability of producing one correct digit is p(s), where s is a scale coordinate such as training loss, compute or effective data. The function can improve smoothly. Exact-match accuracy for an L-digit answer behaves approximately as p(s) raised to L when digit errors are independent. Raising a number smaller than one to the eighth or twentieth power hides early improvement and concentrates visible gains near the top of the range.
This is not a defect in exact match. A bank transfer reference, compiler output or arithmetic answer may need every symbol correct. Exact match answers an important operational question: did the whole artefact pass? It becomes misleading only when the operational score is treated as a direct meter of latent capability.
The same issue appears in multiple-choice grading. Imagine that the logit assigned to the correct answer rises smoothly from just below the leading distractor to just above it. Winner-take-all accuracy jumps from zero to one at the crossing. Brier score, log score and the correct-option margin register the approach before the winner changes. Again, both readings can be useful. They describe different properties.
A metric can be valid for release and invalid for mechanism discovery at the same time. The repair is not to discard hard pass conditions. It is to pair them with rulers that reveal how the system approaches, crosses and behaves beyond the boundary.
The most influential early catalogue defined large-model emergent abilities as abilities absent in smaller models and present in larger ones, with performance not predictable by extrapolating the smaller models. The TMLR survey of emergent abilities made the phenomenon visible across model families and tasks. It also created a precise target for criticism.
Schaeffer, Miranda and Koyejo later showed that many reported cliffs were attached to particular metrics rather than task-model pairs. In their NeurIPS metric analysis, more than 92 per cent of hand-annotated BIG-Bench emergence cases appeared under multiple-choice grade or exact string match. Continuous alternatives such as Brier score or token edit distance exposed graded improvement, and deliberately thresholded vision metrics could manufacture new-looking emergent abilities.
That result should neither be diluted nor universalised. It demonstrates a strong sufficient explanation for many cliffs. It does not prove that every capability changes smoothly under every useful coordinate. The paper itself leaves room for real emergence and analyses fixed outputs within model families. Training dynamics, random-seed mixtures and internal circuit changes remain open.
There are also two different ways for a smooth precursor to meet a discontinuous world. In the first, only the observer discretises: a probability of 0.49 and 0.51 becomes wrong then right under argmax. In the second, the environment itself contains a threshold. A compiler accepts or rejects a program, a tool either has a required permission, and a multi-step plan succeeds only if every dependency resolves. The system may improve smoothly while its consequences change sharply because the environment is conjunctive or irreversible.
That second case is still not evidence of a new internal mechanism, but it can create a real change in risk. If each of ten safety barriers fails with a small probability, correlated improvement or degradation can move the probability of joint failure nonlinearly. A deployment team should therefore keep the operational cliff while refusing to confuse it with a natural boundary in cognition. One curve governs product readiness; another supports scientific explanation.
Sparse observation can make either curve look more mysterious. Three checkpoints below a transition and one above it cannot distinguish a sigmoid, a kink, a discontinuity or a mixture of training outcomes. Logarithmic scale axes can visually compress large intervals, while aggregate means can place the apparent onset between points at which no individual seed changed. A claim of unpredictability needs an explicit forecasting exercise using only earlier checkpoints, not a retrospective impression from the completed curve.
| Ruler | What it rewards | Why a cliff can appear | Best companion measure |
|---|---|---|---|
| Exact string match | Entire sequence correctness | One error makes the score zero | Token edit distance and sequence log probability |
| Multiple-choice accuracy | Winning option | Smooth logit crossing becomes a discrete flip | Brier score, log score and option margin |
| Pass@1 | First sampled solution succeeds | Rare success is poorly resolved in small samples | Pass@k, estimated success probability and uncertainty |
| Aggregate benchmark mean | Average across heterogeneous items | Offset, ceiling and mixture effects hide item curves | Difficulty-conditioned response curves |
| Product acceptance gate | Satisfies all constraints | Conjunction multiplies failure probabilities | Per-constraint risk and joint-failure model |
Thought experiment: the ruler exchange
Imagine two sealed laboratories receive exactly the same model outputs. Laboratory A is told to score exact answers. Laboratory B is told to score normalised edit distance. A reports a capability onset at scale eight. B reports a smooth learning curve from scale two. Neither laboratory can inspect weights or run new generations.
Now exchange their rulers. Their conclusions exchange too. Nothing inside a model changed while the claims changed. This is a decisive diagnosis for M0 sharpness. It does not decide whether a later M2 or M3 transition exists, because neither laboratory has measured internal dynamics or intervened on a mechanism.
The arithmetic is reproducible. The executable example below uses a smooth logistic token probability, then reads it through exact-match, expected edit accuracy and a winner-take-all multiple-choice gate.
from math import exp
def logistic(x: float) -> float:
return 1.0 / (1.0 + exp(-x))
def rulers(scale: float, length: int = 8) -> dict[str, float]:
token_p = logistic(0.75 * (scale - 5.0))
correct_p = token_p
distractor_p = 1.0 - correct_p
return {
"token_probability": token_p,
"expected_token_accuracy": token_p,
"exact_match_probability": pow(token_p, length),
"multiple_choice_grade": float(correct_p > distractor_p),
"two_class_brier": pow(correct_p - 1.0, 2) + pow(distractor_p, 2),
}
rows = [rulers(scale) for scale in range(1, 10)]
assert all(rows[i]["token_probability"] < rows[i + 1]["token_probability"]
for i in range(len(rows) - 1))
assert rows[4]["multiple_choice_grade"] == 0.0
assert rows[5]["multiple_choice_grade"] == 1.0
assert rows[2]["exact_match_probability"] < 0.001
assert rows[-1]["exact_match_probability"] > 0.65
assert all(rows[i]["two_class_brier"] > rows[i + 1]["two_class_brier"]
for i in range(len(rows) - 1))
Mathematical depth: conjunctions sharpen without a phase transition
For a sequence of length L, exact-match probability is p^L under an independence approximation. Its derivative with respect to competence is L × p^(L-1). When p is modest, the derivative is tiny for long strings. Near one, it grows quickly. The metric therefore compresses early progress and expands late progress even when p itself follows a smooth curve.
Dependencies among token errors change the exact formula. They do not remove the general point. Any conjunctive score that requires many conditions to pass can create a steep response from smoother component probabilities. The correct analysis reports both the joint outcome and its component structure, with uncertainty and dependency estimates.
Part III. What survives a ruler change
Changing the metric is the first audit, not the last. A curve can acquire a false cliff through sparse scale points, small test sets, mixed item difficulty, prompt format, decoding policy, contamination, ceiling effects or averages across random seeds. Conversely, averaging can erase a real mixture transition. A smooth mean may combine two sharply different learned strategies.
The unit of analysis must therefore expand from model size to a measurement record:
R = (F, D, T, M, S, P, G, H)
Here F is the model family, D the training-data regime, T the task distribution, M the metric, S the sampling design, P the prompting and decoding policy, G the scale coordinate and H the intervention history. “Ability emerged at 10 billion parameters” suppresses nearly every term.
Model size is often a poor clock for capability. Training loss, effective compute, data quality and architecture can place different models at different functional stages despite similar parameter counts. A NeurIPS study of emergent abilities from the loss perspective reports that models with the same pretraining loss can show similar downstream performance under controlled corpus, tokenisation and architecture conditions, and also finds some task thresholds indexed by loss. That is evidence against a universal metric-mirage account and against parameter count as a sufficient coordinate.
Fine measurement can also expose progress below an apparent floor. The peer-reviewed ICLR PassUntil evaluation uses extensive sampling to give very small success probabilities measurable resolution. It finds task scaling that conventional pass rates miss and reports both predictable scaling and cases of accelerated improvement. The result does not restore every old cliff. It shows why “zero” can mean “below the test’s resolution”.
These results suggest a useful causal chain for an audit. Training changes a distribution over internal states and output probabilities. Decoding converts that distribution into sampled artefacts. A task parser decides what counts as a valid response. A metric maps valid responses into scores. Aggregation combines items, prompts, seeds and models. Plotting then selects axes, smoothing and scale. Every stage can introduce a threshold, and several thresholds can align.
The chain localises responsibility. If re-scoring fixed outputs removes the cliff, training and decoding are unchanged, so the sharpness entered at parsing, metric or aggregation. If the cliff survives re-scoring but moves under decoding temperature, the output distribution may be smooth while the sampling policy exposes it nonlinearly. If it survives ruler and decoding changes yet separates by training seed, the candidate explanation moves upstream toward learned strategy. If a targeted intervention abolishes the high-capability regime, the case for a mechanism becomes substantially stronger.
Forecasting should follow the same chain. Fit a continuous precursor on early checkpoints, propagate uncertainty through the registered operational metric and predict a distribution of possible crossing points. Then compare the observed crossing with that forecast. A capability can cross abruptly in product terms and still be forecastable from smooth precursors. Conversely, a continuous score may depart unexpectedly from its extrapolation without ever producing a visible step. Sharpness and unpredictability are independent properties and should be reported separately.
Worked example: four arithmetic releases
Consider four models tested on 1,000 eight-digit addition items with five prompts and ten random seeds. The figures below are illustrative audit fixtures, chosen to make the method inspectable rather than to report an external benchmark.
| Release | Mean digit accuracy | Exact match | Sequence log score | Seeds using compact routine | Targeted ablation δ exact match (matched random δ) |
|---|---|---|---|---|---|
| A | 0.71 | 0.06 | -4.82 | 0/10 | -0.01 (-0.01) |
| B | 0.79 | 0.15 | -3.61 | 1/10 | -0.03 (-0.02) |
| C | 0.88 | 0.36 | -2.31 | 6/10 | -0.24 (-0.03) |
| D | 0.94 | 0.61 | -1.27 | 9/10 | -0.29 (-0.03) |
Exact match appears to accelerate from B to C. Digit accuracy and log score reveal earlier progress. Yet the seed-level strategy analysis adds information that metric replacement cannot explain: a compact routine becomes common, and its targeted ablation lowers exact match by 24 to 29 percentage points where matched random-state ablation lowers it by 3 points. The careful conclusion has two clauses. Part of the performance cliff is conjunctive scoring. A separate M3 mechanistic reorganisation is compatible with the controlled intervention and deserves replication across new seeds.
The design also prevents a common inference error. C is not a single deterministic phase point. It is a mixture of training outcomes. Saying “the model emerges at C” erases seed variance. The registered object is a distribution over trained systems under a specified recipe.
| Survival test | Hold fixed | Change | Evidence strengthened when |
|---|---|---|---|
| Metric exchange | Outputs | Exact, continuous, calibrated and decomposed scores | Sharpness remains under several justified rulers |
| Resolution increase | Task and model | More items, samples and scale points | The transition narrows rather than dissolves |
| Seed expansion | Training recipe | Random initialisation and data order | A regime boundary appears in the outcome distribution |
| Coordinate exchange | Models | Parameters, compute, loss and effective data | The boundary aligns with a mechanistically meaningful coordinate |
| Intervention test | Baseline capability | Ablate or perturb candidate mechanism | The claimed regime depends selectively on the mechanism |
| Path reversal | Endpoints | Sweep control variable upward and downward | Hysteresis or bistability appears reproducibly |
No single survival test certifies emergence. Together they locate where the discontinuity lives. Metric exchange tests the ruler. Seed expansion tests the training distribution. Intervention tests the mechanism. Path reversal tests dynamical regime structure.
Evaluation depth: prompts, seeds and task mixtures
Prompting can shift a visible onset by supplying decomposition, demonstrations or output constraints. That is not necessarily cheating. It changes the system being evaluated from a bare model to a model-policy pair. The paper should name the pair and avoid comparing it with a different pair under one capability label.
Task mixtures can flatten or sharpen curves. Easy items may reach ceiling while difficult items remain at chance, causing the aggregate to stall. Later gains on difficult items then look sudden. Item-response models and difficulty bins expose this composition. Random seeds create another mixture: averages can place no actual trained model near the mean. Release claims should therefore include per-seed distributions and the probability of entering each regime.
Part IV. Brains, ignition and the ontological surcharge
Consciousness research supplies the hardest version of the ruler problem because the target is contested. A visual stimulus can vary continuously in contrast. Neural responses can be measured through spikes, local field potentials, scalp signals, imaging and report. Each channel has its own resolution and threshold. A participant can use a categorical response for an experience whose confidence or clarity is graded. A global broadcast can be abrupt while the local recurrent processing that feeds it is continuous.
Global neuronal workspace accounts use the idea of ignition: sufficiently strong information gains wide availability across a distributed network. Reviews of conscious processing and workspace dynamics motivate an all-or-none access transition, while empirical work also asks whether awareness is categorical or graded. A multisensory perceptual-awareness study found that the answer depends on what aspect and modality is measured. “Ignition” is therefore a mechanistic hypothesis with operational markers, not a synonym for the appearance of phenomenality.
A threshold in report can mark access, decision or motor commitment even when it is correlated with experience. No-report designs, graded confidence, trial-level neural dynamics and causal perturbation help separate these stages. They do not provide a view from outside every theory.
Phase-transition language can be more than metaphor when it comes with the mathematics and controls of dynamical systems. An order parameter changes across a control variable. Near a continuous transition, correlation length, susceptibility or variance may increase. A discontinuous transition can show jumps, bistability and hysteresis. Finite systems smear ideal singularities, so dense sampling and finite-size analysis matter.
Peer-reviewed models of stochastic spiking networks demonstrate both continuous and discontinuous transitions even when individual neurons have smooth firing probabilities. This is the important counterweight to the metric-mirage result: smooth units can collectively enter distinct regimes. Human brain recordings have also been analysed as occupying a continuum from second-order to first-order critical-like dynamics, including bistability-related signatures. These findings concern dynamical organisation. Their relationship to consciousness requires additional theory and contrastive evidence.
| Proposed signature | What to measure | Stronger interpretation requires | Common confound |
|---|---|---|---|
| Sharp order-parameter change | Registered macro-variable over dense control sweep | Replication and finite-size analysis | Sparse sampling or transformed axis |
| Critical slowing or rising variance | Recovery time and fluctuation statistics | Mechanistic link to the candidate transition | Non-stationarity or measurement noise |
| Bistability | Two stable regimes under matched controls | State-dependent perturbation and dwell-time analysis | Seed mixture treated as within-system dynamics |
| Hysteresis | Different forward and reverse transition points | Controlled path reversal | Irreversible training history |
| Causal macro-variable | Intervention or predictive value beyond selected parts | Robustness across coarse-grainings | Chosen variables encode the answer |
| Conscious access contrast | Report, no-report and neural measures | Theory-specific prediction and causal test | Decision and motor thresholds |
Criticality claims require particular restraint because power laws and avalanche-like distributions can arise from multiple mechanisms, finite samples and analysis choices. Evidence becomes more persuasive when several signatures converge under a generative model, when the proposed control parameter can be manipulated, and when alternative non-critical processes are compared directly. A straight line on log-log axes is an invitation to model comparison, not a phase-transition verdict.
The distinction between a phase change and a functionally important threshold also matters in brains. An organism can exploit a steep but continuous gain curve as if it were categorical. Recurrent amplification can create reliable access without an ideal thermodynamic singularity. Finite neural systems need not reproduce an infinite-system phase transition exactly for the mechanism to be useful. The scientific claim should match the finite system actually measured.
Likewise, a macro-variable can be causally informative without becoming an independent substance. If whole-network synchrony predicts recovery after perturbation better than any chosen neuron, that supports a higher-level description relative to those variables. It does not show that microphysics is incomplete. Causal-emergence measures are especially sensitive to the selected grain, the completeness of the component set and the time horizon. A robust study repeats the analysis across plausible partitions rather than celebrating the partition that maximises novelty.
For machine systems, the analogous test replaces neurons with components, activations, tokens or modules. A collective feature such as a persistent plan may predict later tool use beyond any one local state. The result can justify an agent-level control variable. It still leaves open whether the feature is implemented through distributed computation, external memory, orchestration or a measurement artefact. Emergence language should sharpen that investigation instead of ending it.
Does consciousness emerge?
A physicalist emergentist may say that experience depends on sufficiently organised matter and becomes a real higher-level property when that organisation is achieved. A strong emergentist may add novel powers or laws. A functionalist can locate the relevant transition in causal organisation. An illusionist can explain why a system represents itself as having ineffable properties without adding phenomenal properties to the ontology.
A consciousness-primary orientation reverses the verb. Awareness is not manufactured by complexity; organised systems delimit, express or localise a perspective within awareness. Śaṅkara’s non-dual tradition, surveyed in the Stanford Encyclopedia entry on Śaṅkara, treats consciousness as irreducible rather than as a late product. Process traditions in the West, represented broadly in process philosophy, replace inert substances with becoming and relation. These views do not generate a benchmark prediction merely by being named.
A consciousness-primary view removes the obligation to explain awareness from non-awareness, but it inherits a manifestation problem. Why does this organisation support a bounded perspective, memory continuity, report or suffering while another does not? Which changes alter the contents or boundary of a subject? A manifestation account must still expose conditions and contrasts. Otherwise it converts an explanatory gap into an unmeasured permission.
This is where the ruler audit remains useful across metaphysics. The physicalist asks whether an organisation produces experience. The consciousness-primary researcher asks whether it configures a locus of manifestation. Both should resist treating a benchmark cliff, an activation threshold or a verbal report as a complete answer.
Philosophical depth: emergence, manifestation and category mistakes
Aristotelian form and matter offer one historical way to describe a whole whose capacities are not a mere sum of detached parts. Modern weak emergence can preserve higher-level autonomy within physical determination. Strong emergence pays a larger price by challenging some version of causal closure. Non-dual and process-oriented traditions alter the background ontology more radically.
The traditions should not be collapsed. Saying that awareness is primary differs from saying that every organised system is a conscious subject. Saying that a property supervenes on a base differs from explaining why it exists. Saying that a macro-variable has unique predictive information differs from giving it fundamental causal power. Each move changes the burden rather than dissolving it.
Part V. The ruler-change protocol
An emergence result should ship with an audit record, not a dramatic curve alone. The protocol below is designed for model scaling, agent capability, neural access and other complex systems. It cannot force all domains into one statistic. It forces the claim to disclose which part of the result belongs to the system, the task and the observer.
First, register the claim class or classes. M0 asks whether a score jumps. M1 asks whether an ability becomes operationally usable. M2 asks whether a dynamical regime changes. M3 asks whether a new mechanism carries the capability. M4 asks for ontological autonomy. A study may support several classes, but evidence for one does not automatically count as evidence for another.
Second, freeze the raw record. Preserve item-level outputs, probabilities where available, sampling settings, prompts, seeds, checkpoints and failures. A chart is not raw evidence. Without the record, another evaluator cannot exchange rulers or estimate what the published aggregate concealed.
Third, build a ruler panel before looking for a cliff. Include the metric used for real-world acceptance, a continuous measure of proximity, a calibrated probabilistic score where the task permits it and a decomposition by item difficulty or subskill. Explain why every ruler is connected to the construct.
Fourth, increase resolution on both axes. Add scale points around the suspected transition and enough task samples to resolve rare success. Report uncertainty. A narrow confidence interval around zero is different from a zero caused by forty test items and one decoding attempt.
Fifth, separate population mixtures. Plot each training seed, architecture, prompt policy and meaningful task stratum. Use a hierarchical model when the design warrants one. Look for bimodality and strategy clusters rather than treating every mean curve as one representative system.
Sixth, identify an order parameter only if a dynamical claim is intended. The variable should distinguish regimes before it is fitted to the desired story. Test recovery, variance, state dwell time, forward and reverse paths, and sensitivity to finite system size.
Seventh, intervene on the proposed mechanism. Activation patching, ablation, state reset, pathway interruption, controlled noise or causal stimulation can distinguish a correlated signature from a carrier. Match the intervention to the grain of the claim.
Eighth, write the narrowest statement that survives. “Exact-match performance crosses our acceptance threshold” can be completely correct. “A compact arithmetic routine becomes more common across seeds and is causally necessary under this intervention” is stronger and different. “Reasoning emerged” may be too vague to evaluate.
Ninth, connect the result to a decision. A new operational onset may require access control, red-team coverage or revised forecasts even when it is metric-generated. A mechanistic transition may justify targeted monitoring or a change in training. A consciousness-related marker may justify precaution without licensing certainty about sentience. The decision and the ontology should be documented on separate lines.
Tenth, preserve negative outcomes. If a cliff dissolves under continuous scoring, publish the dissolution. If an apparent order parameter fails path reversal, retain that failure. If an ablation damages every task equally, it did not isolate the proposed mechanism. A research programme learns more from a well-localised non-result than from an emergence label that survives by changing its meaning.
For frontier-model governance, this protocol supports earlier warning. Continuous precursors can be monitored before a hard capability gate is crossed. Seed distributions can expose a rare dangerous strategy before the mean moves. Mechanistic probes can distinguish fluent imitation from a stable capability. None of these measures guarantees that the next training run will remain inside the observed family, but together they replace surprise as the default planning assumption.
The protocol does not ban the word emergence. It makes the word earn a stable referent.
Language that preserves the evidence
Use M0 language when the result is metric-bound: “Exact-match accuracy rises sharply between the sampled checkpoints; token-level scores remain smooth.” Use M1 language for deployment: “The system crosses the registered acceptance boundary under this task distribution.” Use M2 language only after regime evidence: “The order parameter shows a replicated discontinuity with bistability under the registered sweep.” Use M3 language after intervention: “A newly prevalent circuit is necessary for the performance profile under this ablation.”
M4 language should state the metaphysical premise, the dependence relation and the causal claim. No benchmark can silently carry that surcharge. For consciousness, the conclusion must also separate evidence of report, access, integration, subject-boundary conditions and phenomenal presence.
The peer-reviewed TMLR BIG-Bench collaboration was valuable partly because it widened the task surface on which scale effects could be observed. Its descendants should widen the measurement surface too. The same principle applies to evaluations of agents. A binary “completed mission” score may hide smoothly improving planning, tool selection and recovery, while an operational system may still require the binary mission outcome. Preserve both.
Grokking offers another caution. In the original arXiv grokking preprint, generalisation can improve long after training performance saturates. That looks like delayed emergence, but the explanation depends on training dynamics, representation and regularisation rather than model size alone. It is a candidate M2 or M3 phenomenon only after the relevant state variables and mechanisms are tested. The name does not supply the phase theory.
Glossary
| Term | Working meaning in this paper |
|---|---|
| Ruler | The metric, sampling design, resolution and transformation used to turn a system record into a reported result |
| Metric cliff | A sharp change in a reported score that may arise from the measurement operator |
| Operational onset | The point at which a capability satisfies a registered use condition |
| Order parameter | A macro-variable chosen to distinguish dynamical regimes |
| Bifurcation | A qualitative change in system dynamics as a control parameter varies |
| Hysteresis | Path dependence in which forward and reverse sweeps change regime at different points |
| Mechanistic reorganisation | A causally relevant change in the strategy, circuit or process supporting behaviour |
| Causal emergence | Higher-order predictive or causal information not captured by selected components at a declared grain |
| Weak emergence | Higher-level novelty or autonomy compatible with determination by a lower-level basis |
| Strong emergence | A more radical claim of novel powers, laws or efficacy not exhausted by the physical basis |
The decision this changes
Scaling decisions should no longer ask only, “Did a capability emerge?” They should ask four narrower questions. Did a registered score cross a threshold? Did the underlying output distribution change smoothly or discontinuously? Did the system enter a distinct dynamical or mechanistic regime? Which ontological conclusion, if any, is being added to those findings?
For production AI, this changes forecasting and control. Teams should monitor continuous precursor measures even when release depends on binary success. They should estimate the probability of rare capability across seeds and samples. They should not assume that a currently absent hard-pass score has no precursor, nor that every sharp score implies an unforecastable internal leap. Risk controls can be conservative without being conceptually careless.
For consciousness science, the protocol prevents three substitutions. Report is not silently substituted for experience. Neural ignition is not silently substituted for phenomenal onset. Organisational complexity is not silently substituted for ontological production. A physicalist, emergentist and consciousness-primary researcher can share the same audit record while stating different premise bridges.
The publishable result is the strongest invariant that remains after the ruler, resolution and intervention have changed. Sometimes that invariant will be a smooth capability curve with a real operational threshold. Sometimes it will be a reproducible phase transition or new mechanism. Sometimes the evidence will stop before ontology. That stopping point is a result, not a failure of imagination.
The word emergence becomes useful again when it no longer does the work of the missing experiment.