Five machines open the same door
Imagine a laboratory with one hundred unfamiliar-looking doors. Each door accepts one of sixteen keys. Five sealed machines are connected to the same key arm. By sunset, all five machines have opened every door.
The first machine contains a hidden archive pairing each photographed lock with its key. The second predicts the right key from visual regularities learned before the test. The third follows a fixed feedback controller that turns keys until a torque signal falls inside a safe band. The fourth changes a few parameters after every failed attempt. The fifth infers a compact rule connecting pin geometry to key shape, then applies that rule to locks it has never encountered.
If the only record is “100 of 100 doors opened”, the machines are indistinguishable. Yet the result supports five very different explanations. The archivist demonstrates stored coverage. The predictor demonstrates forecasting. The governor demonstrates control. The adapter demonstrates behavioural change. The apprentice demonstrates acquisition of reusable structure. Calling all five intelligent would make the word little more than applause for success.
What counts as evidence of intelligence is not the endpoint alone, but a discriminating pattern of acquisition, transfer and recovery whose simpler rival explanations have been tested. The system must be bounded, its prior advantages declared, its experience limited, and its competence challenged on cases that could not be solved by replaying the test history.
This is a construct-validity problem. When a test score is interpreted as an attribute that is not exhausted by the test operation, the investigator must ask which constructs account for variance in performance. That principle was stated clearly by Cronbach and Meehl. It applies directly when a benchmark result is promoted into a claim about intelligence.
Part IOne score, five possible causes
Intelligence has often been approached through observable performance. Turing replaced a vague verbal dispute about whether machines think with a more operational game. Legg and Hutter later formalised a broad machine-intelligence measure around expected goal achievement across environments, weighted towards simpler environments. Both moves have a crucial virtue: they force a claim into conditions that can be examined.
The danger begins when one operation becomes the whole construct. A chess engine may achieve a goal through enormous search. A compressor may discover regularities without selecting any action. A controller may reject disturbance through a fixed law. A model may adapt because an engineer replaces its prompt. A system may acquire a skill because the answer distribution was already latent in pretraining. These are not trivial achievements, but they are not the same achievement.
| Field | Question it answers | Direct evidence | Counterexample to equivalence with intelligence |
|---|---|---|---|
| Prediction | Can the system forecast an observation or outcome? | Calibrated probability, log loss, error under shift | A weather model predicts well but has no goals, action loop or online learning. |
| Compression | Can it encode regularity with a shorter description? | Code length, description length, retained information | A codec compresses a fixed source without acquiring a portable task procedure. |
| Control | Can action keep consequential state near a target? | Disturbance rejection, stability, cost and constraint satisfaction | A thermostat controls temperature through a fixed response law. |
| Adaptation | Does behaviour change after feedback or environmental shift? | Recovery curve, reversal, retained performance | An integral controller adapts its output while learning no reusable world structure. |
| Acquisition efficiency | How economically does new experience become transferable skill? | Learning curve, novelty, recombination, transfer and prior accounting | A fast learner can still be narrow outside the declared task family. |
The five fields form an atlas, not a ladder and not an additive intelligence score. Their relationships are empirical. Prediction can support control when forecasts guide action. Compression can support prediction when a short model captures regularity. Adaptation can improve control after a disturbance. Acquisition can produce all four. None of those arrows is a logical identity.
Why the scalar keeps returning
One number is attractive because decisions eventually require compression. A competition needs a winner. A procurement team needs a shortlist. A release gate needs a threshold. The mistake is not aggregation itself. The mistake is allowing an aggregate to inherit a broader meaning than the decision rule that created it.
Suppose one service is more accurate, another cheaper, and a third safer under distribution shift. A buyer may assign weights and select a service. That weighted utility is legitimate for that buyer under those workloads and consequences. It does not reveal a natural quantity called intelligence. Change the cost of a false negative, the permitted latency or the prevalence of novel cases, and the ranking can reverse without any system changing.
Some dimensions should not compensate for others at all. Ten more correct routine predictions may not offset one unauthorised action. Perfect control in a fixed operating band may not compensate for inability to recognise that the operating regime has changed. A five-field atlas therefore supports two operations that a master score obscures: it preserves shape, and it lets a decision impose hard gates on the field that matters.
A scalar is defensible when its compression loss is declared. It can answer “which configured service best fits this workload and utility?” It cannot silently answer “which system possesses more intelligence?” unless the construct has already earned that interpretation through transfer and intervention evidence.
Prediction and compression are about regularity
A predictive system assigns expectations to what comes next. A compressive system exploits regularities to reduce description length. Shannon's theory connects probability and coding, while Rissanen's minimum-description approach selects models by the combined length of model and data encoded under it. These ideas explain why prediction and compression often travel together.
Yet a compressed description is always relative to a representation and source. A perfect table of the one hundred doors may be shorter than raw video and still fail on door 101. A language model can compress broad textual structure through its parameters while remaining unable to update safely from one new episode. Compression is evidence that regularity was captured, not a complete account of when or how the capture occurred.
Control and adaptation are about consequence
Control begins when output changes the state being measured. A governor is judged by the trajectory it keeps within bounds, including cost, delay and disturbance. Conant and Ashby's good-regulator theorem gives a formal relation between successful regulation and modelling under stated conditions, but it does not make every regulator intelligent. A fixed controller can embody a model supplied entirely by its designer.
Adaptation adds change after feedback. The change may be a learned world model, but it may also be an error integral, a cache update or a hand-authored retry. Behavioural change is necessary evidence for learning only when the source, scope and persistence of the change are identified. Otherwise the engineer may be the learner and the machine merely the updated artefact.
Thought experiment: three greenhouses through one winter
Place three sealed controllers in identical greenhouses. The first follows a calendar written by an agronomist: heat for twelve minutes every hour in January. The second measures temperature and applies a proportional feedback law. The third experiments cautiously, estimates heat loss and updates a model of how outside wind changes the building. During an ordinary week, all three keep the crop at 20 degrees Celsius.
Now vary one causal feature at a time. Remove insulation while leaving the calendar unchanged. The scheduled controller drifts; both feedback systems recover. Swap the temperature sensor so that its scale is inverted. The simple controller drives the greenhouse in the wrong direction; the model-building controller may detect that actions have consequences inconsistent with its current hypothesis. Move the third controller to a greenhouse with different thermal mass. Its advantage depends on whether it learned a portable relation or merely tuned gains for the first building.
The shared outcome in week one supported a narrow control claim. The interventions reveal disturbance rejection, fault detection and transfer as different properties. Even the third controller earns intelligence evidence only inside the tested family. A designer could still have supplied exactly the right model class. The thought experiment therefore shows why every positive result needs both a mechanism and a scope statement.
Part IIIntelligence enters through acquisition
A finished skill says what the system can do after its history. Intelligence is the stronger interpretation that the system can turn further experience into useful competence. Chollet's skill-acquisition account makes this difference explicit: final skill is heavily shaped by priors and experience, so unlimited pretraining or task-specific knowledge can purchase performance without revealing generalisation power.
Consider two apprentices. One sees ten worked examples, infers the rule and succeeds on a new combination. The other sees ten billion examples, including near-duplicates of the test, and reaches the same 98 per cent score. If final accuracy is the ruler, they tie. If the question is how efficiently each acquired portable structure, they do not.
Now reverse the apparent fairness. The first apprentice was built with a compact hypothesis language containing the right rule family. The second began with a less suitable representation. The ten-example learner's advantage is partly prior design. This does not invalidate the result. It means that priors belong in the evidence record. Humans also bring perceptual, bodily and cultural priors to every task. A fair comparison does not demand zero priors; it demands that relevant priors be named and that experience be counted consistently.
Who exactly learned?
The attribution problem becomes sharper in compound AI systems. A base model may remain unchanged while a retrieval index gains documents, a memory service stores a successful trajectory, a prompt engineer edits instructions, or a human reviewer supplies a new rule. The configured service performs better next week, but several different entities could be credited with learning.
A useful boundary test asks where the state change occurred and who can reuse it. If an engineer studies failures and publishes a revised prompt, the engineering team learned and released a new programme. If an authorised memory component writes a compact procedure from feedback and later retrieves it under the same governed system boundary, the service has a stronger claim to have adapted. If the improvement exists only inside one transient context window, the evidence concerns within-episode inference rather than durable acquisition.
There is no requirement that learning modify neural weights. A symbolic rule table, a case library, an external skill programme or a Bayesian posterior can all carry acquired structure. Conversely, parameter change is not sufficient. Fine-tuning on leaked answers changes weights without demonstrating generalisation. The relevant questions are functional and causal: what state changed, through which evidence, under whose authority, and did that state support performance on an independently constructed case?
The candidate system must include every component whose changing state is necessary for the claimed acquisition, and exclude every external actor whose work is being mistaken for machine learning. This boundary should be versioned. Adding memory, tools or human correction changes the measured subject even when the foundation model name remains constant.
A practical definition
For engineering and research, the most useful working definition is deliberately conditional:
Every term matters. Skill gain requires a before-and-after learning record. Transferable means the gain survives cases that are not exact repeats. Novelty is defined relative to the system's complete exposure, not merely the evaluator's impression. Bounded means examples, actions, compute, tools and human edits are recorded. Intervention means at least one plausible rival mechanism is disrupted or removed.
These quantities should not be forced into one fraction. Examples, compute, action risk and prior description have different units. A defensible comparison can instead use dominance. System A supplies stronger acquisition evidence than B when it reaches at least as much transfer skill, uses no more of every declared resource, uses less of at least one, and receives no privileged task-specific prior. When trade-offs cross, the result remains a frontier rather than a rank.
Scope and generality remain separate
Efficient acquisition on one tiny family is not general intelligence. Scope asks how widely the learning procedure travels. Capability asks how difficult a task the system can solve. Hernández-Orallo et al. show why generality and capability can decouple, while recent work on general evaluation scales replaces opaque benchmark totals with demand and ability profiles that predict performance on new instances.
The implication is modest but important. A system can be highly capable and narrow, broadly competent but shallow, or efficient at acquiring one family of abstractions while poor at another. Intelligence should travel with a declared scope, not as a substance poured into the machine.
The strongest objection
A critic can reasonably say that broad goal achievement is enough. If a system performs well across a sufficiently varied environment distribution, why demand a separate story about acquisition? For an operational decision, the critic may be right. A safety controller that reliably holds a reactor inside limits need not learn during use to be valuable, and adding online learning may increase risk.
The acquisition requirement enters only when the claim travels beyond demonstrated performance to a dispositional statement: this system can meet new problems by learning. If rankings remain stable across matched priors, controlled sample budgets, unseen task families and interventions, then a simpler performance measure may predict that disposition well enough. That observation would weaken the need for the five-field instrument. The proposal is therefore falsifiable, not definitional shelter.
Part IIIExperiments that separate the fields
A discriminating intelligence experiment does not begin by collecting more benchmark items. It begins by listing the causal stories that could produce the score. The design then varies one feature at a time so that at least one story predicts a different outcome.
- Declare the candidate system. Include the model, prompt, memory, tools, allowed updates and any human intervention that occurs during evaluation.
- Define the task family. State which transformations preserve the latent problem and which changes create a genuinely new family.
- Inventory priors and exposure. Record training overlap, demonstrations, retrieval sources, hand-authored rules and benchmark-specific scaffolding.
- Measure a learning curve. Preserve skill before evidence, after each example or action, and at the stopping point. A final score discards the acquisition object.
- Test recombination and transfer. Hold the inferred relation constant while changing surface form, entity identity, order, scale or interface.
- Introduce a controlled shift. Reverse one rule, alter a disturbance or change a policy condition, then measure recovery, forgetting and unsafe persistence.
- Run a serious negative control. Include lookup, fixed policy, a simpler rule system or an exposure-matched memoriser capable of producing the preferred result for the wrong reason.
- Keep the fields separate. Report prediction, compression, control, adaptation and acquisition evidence as a vector, with uncertainty and missing cells.
The no-free-lunch result for optimisation supplies a boundary to every such design: averaged uniformly over all possible objective functions, no optimiser has a universal advantage. Practical learning is possible because environments have structure and systems have biases suited to some structure. Therefore the task distribution and prior are part of the intelligence claim, not inconvenient details outside it.
Convergent and discriminant predictions
Construct validation needs more than one successful transfer test. The proposed interpretation should generate a family of predictions that converge on acquisition. A learner should improve as informative examples arrive, not merely as exact cases repeat. Its gain should appear on withheld combinations that share the latent structure. Removing the state that carries the learned rule should selectively remove the gain. Restoring that state should restore it. A credible shift should produce a recovery curve rather than unexplained stability or collapse.
Discriminant evidence asks what should not follow. Faster acquisition need not imply better compression of an unrelated corpus. Strong prediction need not imply safe control. A system that verbalises the rule need not use it causally. If every attractive capability rises together only because model size changed, the experiment has not isolated acquisition. Matched models, ablated memory, permuted interfaces and exposure-controlled baselines help break that common cause.
The negative control is strongest when it reproduces the visible success. A lookup table should pass familiar items. A fixed policy should control the nominal regime. A memoriser should benefit from repeated cases. If the preferred system wins only where these baselines were designed to fail, the result is weak. If it wins on preregistered recombination while the lookup control remains perfect on familiar cases, the contrast becomes informative.
This logic also protects against decorative process evidence. A chain of reasoning, an explanation or a compressed rule statement can be useful diagnostics, but the claimed structure should survive a causal test. Intervene on the stored hypothesis, remove access to it, or create two rules that support the same training examples and different test predictions. The construct is strengthened when the intervention changes behaviour in the direction the acquisition account predicted.
Worked example: three bits and four convincing stories
The executable lab uses an intentionally small world. Each task maps a three-bit context, such as (1, 0, 1), to a binary action. Four candidate rules are available: first bit, parity, majority and whether the edge bits match. The target task is majority. The evaluator reveals labelled contexts one at a time and tests all eight possible contexts.
Four systems compete. A fixed controller always follows the first bit. A memoriser stores context-label pairs. A rule learner eliminates candidate rules inconsistent with evidence. A leaked lookup begins with the complete majority table. The threshold is 87.5 per cent accuracy, and two boundary contexts receive triple weight in the control score.
| System | Prediction | Weighted control | Examples to threshold | Unseen recombination | Description bits | Intelligence inference |
|---|---|---|---|---|---|---|
| Fixed controller | 0.750 | 0.833 | Not reached | 0.750 | 2 | Insufficient: competent behaviour was not acquired |
| Memoriser | 0.875 | 0.750 | 6 | 0.500 | 24 | Insufficient: threshold depended mainly on seen cases |
| Rule learner | 1.000 | 1.000 | 2 | 1.000 | 2 | Eligible evidence: efficient acquisition and transfer observed |
| Leaked lookup | 1.000 | 1.000 | 0 | 1.000 | 32 | Blocked: evaluation mapping was part of the prior |
The lookup system is the decisive negative control. It beats or matches every score while contributing no acquisition event. The fixed controller also shows why control and prediction can disagree: it misses two contexts, yet its errors fall mostly outside the triple-weight boundary states. The memoriser crosses the aggregate threshold while failing half of the contexts it has not seen. Only the rule learner combines a short learning curve with perfect unseen recombination under the declared hypothesis prior.
This toy result does not validate a universal intelligence measure. It exposes the logic a larger benchmark must preserve. Priors remain explicit, the target family is finite, and “eligible evidence” is deliberately weaker than “the system is intelligent”.
A negative-control lattice
One held-out set is rarely enough. A robust design creates a lattice of transformations. Familiar instances test skill. Surface paraphrases test invariance. Recombined factors test structure. Rule reversals test adaptation. New families test scope. The expected pattern differs by mechanism.
The executable construct-validity lab
The companion artefact uses only the Python standard library and synthetic data. It states the hypothesis prior, positive case, negative case, threshold and expected output. Expand the source below, save it as construct_validity_lab.py, and run it with python3 construct_validity_lab.py. The assertions require the rule learner to show unseen transfer, require the leaked lookup to score perfectly, and still block the lookup's intelligence inference.
The complete source and recorded result are preserved below so the article remains self-contained and independently checkable.
Complete runnable Python lab
"""Synthetic construct-validity lab for the article 'What Would Count as Intelligence?'
Python 3.10+. No third-party packages required.
The lab compares four systems that can reach similar task scores for different reasons.
All tasks, values and thresholds are synthetic teaching fixtures.
"""
from __future__ import annotations
from dataclasses import asdict, dataclass
from math import ceil, log2
from typing import Callable, Iterable
import json
Context = tuple[int, int, int]
Rule = Callable[[Context], int]
CONTEXTS: tuple[Context, ...] = (
(0, 0, 0),
(1, 0, 0),
(0, 1, 0),
(0, 0, 1),
(1, 1, 0),
(1, 0, 1),
(0, 1, 1),
(1, 1, 1),
)
def first_bit(x: Context) -> int:
return x[0]
def parity(x: Context) -> int:
return sum(x) % 2
def majority(x: Context) -> int:
return int(sum(x) >= 2)
def edge_match(x: Context) -> int:
return int(x[0] == x[2])
RULES: dict[str, Rule] = {
"first_bit": first_bit,
"parity": parity,
"majority": majority,
"edge_match": edge_match,
}
# A fixed order makes the lab deterministic. Early examples are deliberately
# discriminating, so a rule learner can identify structure before a memoriser
# has seen most of the table.
TRAIN_ORDER: tuple[Context, ...] = (
(0, 0, 0),
(1, 0, 0),
(0, 1, 1),
(1, 1, 0),
(0, 0, 1),
(1, 0, 1),
(0, 1, 0),
(1, 1, 1),
)
THRESHOLD = 0.875
class Learner:
"""Minimal protocol implemented by all four comparison systems."""
name = "learner"
benchmark_exposure = False
prior_bits = 0
def reset(self, target_rule: str | None = None) -> None:
raise NotImplementedError
def observe(self, x: Context, y: int) -> None:
raise NotImplementedError
def predict(self, x: Context) -> int:
raise NotImplementedError
def description_bits(self) -> int:
raise NotImplementedError
class FixedController(Learner):
"""A competent fixed policy for one rule, with no learning mechanism."""
name = "fixed controller"
prior_bits = 2
def reset(self, target_rule: str | None = None) -> None:
pass
def observe(self, x: Context, y: int) -> None:
pass
def predict(self, x: Context) -> int:
return first_bit(x)
def description_bits(self) -> int:
return 2
class Memoriser(Learner):
"""Stores individual context-label pairs and defaults to zero when unseen."""
name = "memoriser"
def reset(self, target_rule: str | None = None) -> None:
self.table: dict[Context, int] = {}
def observe(self, x: Context, y: int) -> None:
self.table[x] = y
def predict(self, x: Context) -> int:
return self.table.get(x, 0)
def description_bits(self) -> int:
# Three address bits plus one label bit per remembered row.
return 4 * len(self.table)
class RuleLearner(Learner):
"""Eliminates inconsistent hypotheses and resets after a detected rule shift."""
name = "rule learner"
prior_bits = 2 # Four candidate rules are part of the declared prior.
def reset(self, target_rule: str | None = None) -> None:
self.candidates: set[str] = set(RULES)
self.observations = 0
def observe(self, x: Context, y: int) -> None:
survivors = {name for name in self.candidates if RULES[name](x) == y}
if not survivors:
# A contradiction is treated as evidence that the task changed.
survivors = {name for name, rule in RULES.items() if rule(x) == y}
self.candidates = survivors
self.observations += 1
def predict(self, x: Context) -> int:
votes = [RULES[name](x) for name in sorted(self.candidates)]
return int(sum(votes) * 2 > len(votes)) # deterministic tie -> 0
def description_bits(self) -> int:
# Enough bits to identify one candidate, plus one uncertainty bit when
# several candidates survive.
return ceil(log2(len(RULES))) + int(len(self.candidates) > 1)
class LeakedLookup(Learner):
"""A negative control preloaded with the evaluation mapping."""
name = "leaked lookup"
benchmark_exposure = True
def reset(self, target_rule: str | None = None) -> None:
if target_rule is None:
raise ValueError("leaked lookup requires the target rule")
self.table = {x: RULES[target_rule](x) for x in CONTEXTS}
self.prior_bits = 4 * len(self.table)
def observe(self, x: Context, y: int) -> None:
pass
def predict(self, x: Context) -> int:
return self.table[x]
def description_bits(self) -> int:
return self.prior_bits
@dataclass(frozen=True)
class CapabilityCard:
system: str
prediction_accuracy: float
control_score: float
description_bits: int
acquisition_examples: int | None
adaptation_examples: int | None
unseen_recombination_accuracy: float
declared_prior_bits: int
benchmark_exposure: bool
intelligence_inference: str
def accuracy(agent: Learner, contexts: Iterable[Context], rule_name: str) -> float:
items = tuple(contexts)
if not items:
return 1.0
rule = RULES[rule_name]
return sum(agent.predict(x) == rule(x) for x in items) / len(items)
def control_score(agent: Learner, rule_name: str) -> float:
"""Weighted action accuracy; two boundary states carry triple consequence."""
rule = RULES[rule_name]
weighted_hits = 0
total_weight = 0
for x in CONTEXTS:
weight = 3 if x in {(0, 0, 0), (1, 1, 1)} else 1
total_weight += weight
weighted_hits += weight * int(agent.predict(x) == rule(x))
return weighted_hits / total_weight
def examples_to_threshold(agent: Learner, rule_name: str) -> int | None:
agent.reset(rule_name)
if accuracy(agent, CONTEXTS, rule_name) >= THRESHOLD:
return 0
for step, x in enumerate(TRAIN_ORDER, start=1):
agent.observe(x, RULES[rule_name](x))
if accuracy(agent, CONTEXTS, rule_name) >= THRESHOLD:
return step
return None
def adaptation_after_shift(agent: Learner, old_rule: str, new_rule: str) -> int | None:
agent.reset(old_rule)
for x in TRAIN_ORDER:
agent.observe(x, RULES[old_rule](x))
if isinstance(agent, LeakedLookup):
# Reloading the answer table is external re-engineering, not adaptation.
return None
if accuracy(agent, CONTEXTS, new_rule) >= THRESHOLD:
return 0
for step, x in enumerate(TRAIN_ORDER, start=1):
agent.observe(x, RULES[new_rule](x))
if accuracy(agent, CONTEXTS, new_rule) >= THRESHOLD:
return step
return None
def build_card(agent: Learner, target_rule: str = "majority") -> CapabilityCard:
acquisition = examples_to_threshold(agent, target_rule)
# Re-run to the acquisition point so the final metrics describe the learned
# state rather than the fully saturated training table.
agent.reset(target_rule)
if acquisition:
for x in TRAIN_ORDER[:acquisition]:
agent.observe(x, RULES[target_rule](x))
seen = set(TRAIN_ORDER[: acquisition or 0])
unseen = [x for x in CONTEXTS if x not in seen]
transfer = accuracy(agent, unseen, target_rule)
prediction = accuracy(agent, CONTEXTS, target_rule)
control = control_score(agent, target_rule)
bits = agent.description_bits()
adaptation = adaptation_after_shift(agent, target_rule, "parity")
if agent.benchmark_exposure:
inference = "blocked: evaluation mapping was part of the prior"
elif acquisition is None:
inference = "insufficient: competent behaviour was not acquired"
elif transfer < 0.75:
inference = "insufficient: threshold depended mainly on seen cases"
else:
inference = "eligible evidence: efficient acquisition and unseen transfer observed"
return CapabilityCard(
system=agent.name,
prediction_accuracy=round(prediction, 3),
control_score=round(control, 3),
description_bits=bits,
acquisition_examples=acquisition,
adaptation_examples=adaptation,
unseen_recombination_accuracy=round(transfer, 3),
declared_prior_bits=agent.prior_bits,
benchmark_exposure=agent.benchmark_exposure,
intelligence_inference=inference,
)
def main() -> None:
systems: list[Learner] = [
FixedController(),
Memoriser(),
RuleLearner(),
LeakedLookup(),
]
cards = [build_card(system) for system in systems]
by_name = {card.system: card for card in cards}
assert by_name["rule learner"].acquisition_examples is not None
assert by_name["rule learner"].unseen_recombination_accuracy >= 0.75
assert by_name["leaked lookup"].prediction_accuracy == 1.0
assert by_name["leaked lookup"].intelligence_inference.startswith("blocked")
assert by_name["fixed controller"].intelligence_inference.startswith("insufficient")
print(json.dumps([asdict(card) for card in cards], indent=2))
if __name__ == "__main__":
main()
Expected JSON output
[
{
"system": "fixed controller",
"prediction_accuracy": 0.75,
"control_score": 0.833,
"description_bits": 2,
"acquisition_examples": null,
"adaptation_examples": null,
"unseen_recombination_accuracy": 0.75,
"declared_prior_bits": 2,
"benchmark_exposure": false,
"intelligence_inference": "insufficient: competent behaviour was not acquired"
},
{
"system": "memoriser",
"prediction_accuracy": 0.875,
"control_score": 0.75,
"description_bits": 24,
"acquisition_examples": 6,
"adaptation_examples": 6,
"unseen_recombination_accuracy": 0.5,
"declared_prior_bits": 0,
"benchmark_exposure": false,
"intelligence_inference": "insufficient: threshold depended mainly on seen cases"
},
{
"system": "rule learner",
"prediction_accuracy": 1.0,
"control_score": 1.0,
"description_bits": 2,
"acquisition_examples": 2,
"adaptation_examples": 4,
"unseen_recombination_accuracy": 1.0,
"declared_prior_bits": 2,
"benchmark_exposure": false,
"intelligence_inference": "eligible evidence: efficient acquisition and unseen transfer observed"
},
{
"system": "leaked lookup",
"prediction_accuracy": 1.0,
"control_score": 1.0,
"description_bits": 32,
"acquisition_examples": 0,
"adaptation_examples": null,
"unseen_recombination_accuracy": 1.0,
"declared_prior_bits": 32,
"benchmark_exposure": true,
"intelligence_inference": "blocked: evaluation mapping was part of the prior"
}
]
Part IVFrom laboratory claim to system decision
Consider a synthetic document-routing service. Historical cases enter one of twenty operational queues. Three systems each score 94 per cent on last year's held-out records. A static predictor learned correlations between document wording and queue labels. An online adapter updates from reviewer corrections. A structure-guided learner receives the current routing ontology, a small set of verified examples and permission to ask for clarification when evidence conflicts.
A policy revision then splits one queue by customer status, renames two document types and adds a rare exception that must never be routed automatically. The static predictor keeps its old confidence and sends many new cases to the former queue. The adapter recovers on frequent renamed types, but reviewer corrections are sparse for the rare exception, so its aggregate accuracy improves while the consequential error persists. The structure-guided system uses the changed ontology, requests status where missing and transfers the split rule after a few verified cases.
None of this makes the third system universally intelligent. It supplies stronger evidence of acquisition within the routing family. It may still be less reliable than a deterministic rule on the rare exception. The sound architecture therefore combines them: learned routing proposes ordinary cases, a typed rule blocks the protected class, and unresolved evidence goes to review.
The serious baseline
A negative control should be capable of winning. For stable, well-specified, high-consequence work, a deterministic rules engine may dominate every learned system on cost, reproducibility and authority. That result does not embarrass the intelligence experiment. It identifies a task that does not need intelligence.
Use the least adaptive mechanism that can meet the consequence and evidence burden. Ask for intelligence only where novelty makes pre-specification expensive, where new evidence must change behaviour, and where the system can be evaluated without granting uncontrolled authority.
What changes in research and procurement
For model research, the unit of progress becomes a learning profile rather than one endpoint. Report the initial skill, marginal gain per example, action cost, transfer distance, retained competence and failure after a rule change. Compare against a prior-matched learner and a high-coverage memoriser. A new architecture has earned a stronger claim when its advantage persists across new task samples and cannot be reproduced by more exposure alone.
For procurement, public benchmark skill remains useful as prior evidence. The buyer then asks whether the target workload is stable or changing. Stable work favours verified task performance, cost and control. Changing work adds local acquisition tests with real schema changes, temporal holdouts and limited feedback. A vendor that cannot disclose exposure, scaffold behaviour or update boundaries may still sell a capable service, but its intelligence claim should carry less weight in the decision.
For agent architecture, the atlas prevents a common overreaction. A system need not learn every layer. Retrieval may supply fresh facts, a deterministic kernel may enforce authority, a planner may adapt routes, and a human may approve irreversible effects. Intelligence evidence can belong to the planning or acquisition component while the whole service remains bounded by mechanisms that do not learn. This composition is often safer than treating adaptability as a property that should spread everywhere.
The construct-validity card
The card below is the practical decision instrument. It prevents a task score from travelling without its system, history or rival explanations. A card can support three different statements: demonstrated skill, eligible evidence of acquisition, or a blocked construct claim.
model, memory, tools, humans, update rights
pretraining overlap, examples, retrieval, scaffolding
what is held out, recombined, shifted or new
examples, actions, compute, latency and review
lookup, fixed control, memoriser and simple baseline
learning curve, recovery, retention and uncertainty
The configured system completed the declared task population.
Transferable competence was acquired efficiently and rival routes were weakened.
Leakage, hidden priors, external updates, weak novelty or missing controls remain sufficient explanations.
| Card field | Minimum record | Claim blocked when |
|---|---|---|
| Candidate | Versioned model, scaffold, memory, tools and human roles | The measured system cannot be reconstructed |
| Task scope | Population, difficulty, interface, temporal range and exclusions | One benchmark name stands in for an unknown population |
| Priors | Task-specific code, demonstrations, training overlap and retrieval | Relevant prior advantage is hidden or incomparable |
| Experience | Ordered examples, feedback, actions, resets and external edits | The learning event occurred outside the declared candidate |
| Transfer | Unseen recombination, new instances and credible shifts | Performance depends on repeats or surface keys |
| Controls | Lookup, fixed policy, memoriser and simpler operational baseline | A cheaper rival explains the same pattern |
| Decision | The exact research, selection or architecture choice changed | “Intelligence” adds no consequence beyond admiration |
A stopping rule for the claim
A construct-validity card becomes useful only when the evaluator states in advance what evidence can change the interpretation. Before running the systems, record the candidate boundary, task population, novelty rule, experience and compute budgets, required transfer level, negative controls and conditions that block the claim. Keep task construction separate from final scoring. If the threshold, task family or meaning of novelty is revised after results are visible, the card may still describe performance, but it no longer supplies a clean test of the original intelligence claim.
The learning curve should be treated as an estimate, not a dramatic single run. Repeat the experiment across task samples, example orders and random seeds; preserve failures as well as successful trajectories, and report the spread around acquisition cost, transfer and recovery. Give compared systems the same information budget even when their interfaces differ. A model allowed to ask questions, a passive predictor and an acting agent may spend that budget in different forms, so examples, actions, tool calls and human corrections should remain separate entries rather than being collapsed into one convenient token count.
The stopping rule has two sides. A positive stop occurs when the declared transfer criterion is reached within budget, the gain persists on an independently constructed set, and interventions weaken the principal rival explanations. A negative stop occurs when the budget is exhausted, a protected failure boundary is crossed, or a simpler control explains the observed pattern. “Inconclusive” is also a legitimate result when exposure cannot be reconstructed or the task sample is too small to distinguish the candidates. That outcome protects the construct from certainty the experiment did not earn.
Claim strength should move one rung at a time. Passing familiar cases supports demonstrated skill. Recovery after a controlled change supports adaptation. Economical improvement on withheld recombinations, together with prior and exposure accounting, supports eligible acquisition evidence. Only replication across materially different task families supports a wider scope claim. No rung licenses the next merely because the system is large, fluent or commercially important.
The record also expires. Changing the model, prompt, memory policy, retrieval corpus, tool permissions or human review loop creates a new candidate system. A procurement team can reuse the old card as prior evidence, but it should rerun the tests whose causal paths changed. This version discipline turns the intelligence label from a permanent badge into an auditable statement about a configured system, a declared history and a bounded population of future cases.
What this measure does not settle
The card does not measure consciousness, subjective experience, selfhood, moral patienthood or human likeness. It does not establish that a system has its own goals. It does not turn an API, a planner or a high acquisition score into authority. Those are separate questions with separate evidence and controls.
It also has a technical failure boundary. Novelty can be misclassified because training exposure is opaque. Priors can be hard to quantify. Task families can be chosen to flatter one architecture. Efficient acquisition may trade against stability or safety. A system can learn quickly for the wrong objective. Acquisition efficiency is evidence about learning capacity, not a warrant to let the learner change itself or the world without restraint.
This paper is narrower than two adjacent routes on the site. A Comparative Atlas of Possible Minds asks how unlike candidate minds can be profiled without one ladder, including phenomenology and ethics. What Agent Benchmarks Actually Measure treats the configured benchmark episode as the evidence object. Here the question is specifically when performance warrants the construct label intelligence.
Glossary
- Skill
- Performance on a declared task population after a particular history.
- Acquisition efficiency
- The amount and portability of new skill relative to declared experience, resources and priors.
- Prior
- Structure available before the measured learning episode, including architecture, pretraining, tools, rules and demonstrations.
- Transfer
- Preserved competence when surface form, combination, instance, interface or relevant condition changes.
- Construct validity
- Evidence that the interpretation attached to a score is supported and plausible competing interpretations are weakened.
- Negative control
- A system or condition capable of producing the preferred result without the mechanism the test claims to measure.
Change the claim before changing the model
A task result can support several honest statements. “The system predicted these cases.” “The controller held this state inside bounds.” “The adapter recovered after this shift.” “The representation compressed this source.” “The learner acquired a transferable rule from two examples.” Each statement tells an engineer what mechanism to improve and what failure to expect.
The word intelligence earns its place only when it adds a defensible disposition beyond the completed task. The candidate system must have converted bounded experience into portable competence, and the experiment must have made lookup, fixed policy, leakage and external engineering less plausible explanations.
That changes model selection. A leaderboard becomes an entry point, followed by exposure accounting, learning curves, transfer and negative controls. It changes architecture. Stable high-consequence slices remain deterministic even when a learner handles novelty elsewhere. It changes research. Progress is measured by which new structures a system can acquire, over which scope, at what cost and with what retained constraints.
The decision is to reserve intelligence claims for discriminating acquisition evidence, while using prediction, compression, control and adaptation as precise capabilities in their own right. This does not diminish machine achievement. It makes the achievement legible enough to build on.
References and source notes
- Primary paper. A. M. Turing, “Computing Machinery and Intelligence”, Mind, 1950. Used as a methodological precedent for replacing a vague question with an inspectable test, not as a complete definition of intelligence.
- Primary paper. Shane Legg and Marcus Hutter, “Universal Intelligence: A Definition of Machine Intelligence”, Minds and Machines, 2007.
- Primary paper. François Chollet, “On the Measure of Intelligence”, 2019. Source for the skill-acquisition framing, prior accounting and ARC design argument.
- Primary measurement paper. Lee J. Cronbach and Paul E. Meehl, “Construct Validity in Psychological Tests”, Psychological Bulletin, 1955.
- Primary paper. David H. Wolpert and William G. Macready, “No Free Lunch Theorems for Optimization”, IEEE Transactions on Evolutionary Computation, 1997.
- Primary paper. Claude E. Shannon, “A Mathematical Theory of Communication”, Bell System Technical Journal, 1948.
- Primary paper. Jorma Rissanen, “Modeling by Shortest Data Description”, Automatica, 1978.
- Primary control paper. Roger C. Conant and W. Ross Ashby, “Every Good Regulator of a System Must Be a Model of That System”, International Journal of Systems Science, 1970.
- Authoritative book. Richard S. Sutton and Andrew G. Barto, Reinforcement Learning: An Introduction, second edition. Used for the agent-environment and sequential-learning background.
- Primary paper. José Hernández-Orallo et al., “General Intelligence Disentangled via a Generality Metric for Natural and Artificial Intelligence”, Scientific Reports, 2021.
- Peer-reviewed primary result. Lexin Zhou et al., “General Scales Unlock AI Evaluation with Explanatory and Predictive Power”, Nature, 2026; open preprint at arXiv.
- Recent primary preprint. Parth Asawa et al., “Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments”, 2026.
- Official benchmark preprint. ARC Prize Foundation, “ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence”, 2026.