The field that caused the hold
At 09:12, an AI-assisted account-review service begins placing almost every application on hold. The dashboard offers an immediate clue. A field called evidence_confidence falls below 0.45 in every held case and remains above 0.45 in nearly every approved case. The relationship is so clean that the incident channel settles on a sentence within minutes: “low confidence is causing the holds”.
An engineer then changes the value in a staging environment. She pins the field to 0.80 while leaving the incoming applications untouched. The holds disappear. This is not merely another chart. The intervention changes the outcome, so the causal claim has earned support.
Now ask a different question: what constitutes a hold decision? The confidence scorer emits a number. A policy gate interprets that number alongside identity, purpose and account state. An authority service writes the hold. A readback confirms that the state changed. The confidence field can cause the gate to choose the hold path, yet it does not by itself make a governed hold exist. Remove the policy gate and the number has no authorised meaning. Remove the state write and there is a proposal but no operational hold.
Dependency is not a single ladder from weak evidence to strong evidence. It is a family of different questions. The surface asks what travels together. The next depth asks what changes under a controlled difference. The deepest asks what organised activity realises the target phenomenon. Confusing these depths produces some of the most persistent errors in science, engineering and claims about artificial minds.
The trace is not the lever
Begin with the least controversial observation: two measurements differ together. That fact may be valuable. It may support forecasting, triage or anomaly detection. It does not yet tell us which difference produced the other.
An observation belongs to a regime
Suppose application holds occur more often when evidence confidence is low. The statement is incomplete until we name how the cases entered the data. Was confidence measured before the policy gate ran? Did the service request extra evidence only for risky cases? Were manually reviewed cases excluded? Did a threshold compress a continuous score into two bins? The observed relationship belongs to this measurement and selection regime, not to an unlabelled world.
Correlation here means any statistical dependence, not only Pearson’s coefficient. A conditional probability, mutual information estimate, regression coefficient, embedding similarity or benchmark score can all expose dependence. The form changes; the epistemic limit does not. A pattern can be excellent for prediction while remaining radically ambiguous about production.
Imagine three services with exactly the same joint distribution of confidence and holds. In Service A, confidence is read by the policy gate and changes the decision. In Service B, a hidden risk score independently lowers confidence and triggers the hold. In Service C, the hold workflow requests harder documents, and those documents later lower the confidence score. The dashboard is identical. The causal stories are not.
Now intervene on confidence alone. A changes, B does not, and C may change only if the workflow also reads the edited value. The intervention reveals a difference that observation could not.
The smallest useful causal question
A causal question compares outcomes under alternative assignments. In potential-outcome notation, each case has a possible outcome under treatment and under no treatment, although only one is observed. In a structural causal model, the do operator represents replacing the normal assignment rule for a variable with an external setting. These are different formalisms for the same practical shift: stop asking what happened among naturally occurring groups and ask what would differ under a specified change.[2][1]
Interventionist explanation adds a useful discipline: the change must be connected to a counterfactual pattern, not merely followed by an outcome. A proposed cause earns explanatory relevance when suitably changing it would change the outcome under specified background conditions.[4] That still leaves open whether the changed item is a trigger, an external condition or a constituent of the process.
Δcause = E[Y | do(X = 1)] − E[Y | do(X = 0)] Here X is the candidate intervention and Y is the target outcome. The two contrasts coincide only under assumptions that make the observed groups exchangeable for the causal question.
The notation compresses a vital design choice. What exactly is set? For whom? At what time? Compared with what alternative? Over what outcome window? “Does confidence cause holds?” is still too vague. “For applications eligible for automated review, what is the 30-second effect on the probability of an authorised hold when the policy input is set to 0.80 rather than its computed value?” is an estimand.
Randomisation can make assignment independent of many rival causes. Observational identification can also be defensible when a causal graph, design and domain assumptions justify adjustment, instruments, discontinuities or other strategies. The point is not that experiments are the only route. It is that a causal estimate is an observational result plus an identification argument. The assumptions are part of the claim, not a footnote.[3]
A reversal you can calculate mentally
A service enables a cache mainly during high traffic. Under low traffic, uncached requests average 100 ms and cached requests 80 ms. Under high traffic, uncached requests average 240 ms and cached requests 180 ms. The cache helps in both regimes.
Yet 80 per cent of cached observations come from high traffic, while 80 per cent of uncached observations come from low traffic. The pooled averages become 160 ms with cache and 128 ms without it. A dashboard now “shows” that enabling the cache adds 32 ms. Balance the traffic mix or randomise cache assignment, and the estimated effect becomes a 40 ms reduction. The sign reverses because traffic influenced both assignment and latency.
The worked example also exposes a limit. Even after the intervention establishes that enabling the cache reduces latency, it has not explained what latency is, how request serving is organised, or whether the cache is part of the mechanism for this service. Causal success finishes one question and opens another.
A causal estimate has an address
A result such as “X changes Y” is incomplete until its address is supplied. Which population was eligible? Which version of X was set? What was the comparator? Over what time window was Y measured? Which other pathways were held fixed? A causal effect belongs to that package. Moving the number to a new population, interface or system configuration is a transport claim that needs fresh support.
This also separates a population effect from the cause of one episode. An average intervention can change the rate of an outcome while leaving many individual cases unchanged. Conversely, a factor can matter decisively in one case even when its average effect is small because it rarely becomes active. Average causal effect, actual cause and mechanism are three outputs, not synonyms.
Take 400 synthetic applications that already satisfy the same eligibility rules. In staging, assign 200 applications to the normal decision path. Assign the other 200 to an external setter that replaces the computed confidence input with 0.80 immediately before the policy evaluates it. The normal arm produces 62 holds, or 31 per cent. The set-to-0.80 arm produces 38 holds, or 19 per cent. The estimated risk difference is −12 percentage points.
The intervention supports a precise claim: for this eligible staging population and this policy configuration, replacing the naturally assigned confidence field with 0.80 reduced the probability of a hold by about twelve percentage points. Randomisation protects that contrast from many pre-treatment differences between the groups. It does not show that increasing model quality would have the same effect. The setter bypassed the scorer, so the experiment targets the policy input rather than the process that normally generates it.
Now inspect one application held in the normal arm. Replaying it with confidence fixed at 0.80 does not release it because an independent sanctions rule also requires a hold. The field had a population effect but was not the difference-maker in this episode. A second application is released under the same replay, so the field is a plausible actual cause there. The average result could not tell those cases apart.
Finally, suppose the setter also prevents a timeout branch that normally activates while the scorer is running. The observed −12 points now combine at least two pathways: the value delivered to policy and the timing behaviour of the decision service. The experiment remains randomised, yet its treatment is no longer the clean variable named in the headline. Randomisation solves assignment bias. It does not automatically solve treatment ambiguity, interference, non-compliance or off-target effects.
Permitted conclusion: the implemented setter changed hold frequency under the stated configuration. Prohibited conclusion: the scorer is the mechanism of holding, or confidence caused every observed hold. To move further, the team needs path-specific tests, case-level replay and an account of the configured organisation.
The lever is not the machine
A cause answers why an event or difference occurred. A constitutive explanation answers how a capacity, state or process is realised by organised parts and activities. The two often meet inside mechanisms, but they are not interchangeable.
Which “why” are you asking?
Ask why this application was held. A stale retrieval index, a low confidence input and a policy threshold may form an actual causal path. Ask instead how the deployed service has the capacity to place an application on hold. The answer now describes a configured organisation: evidence retrieval, model proposal, policy decision, authority, state transition and readback. The incident cause is one trajectory through that organisation. The constitutive explanation is the organisation that makes such trajectories possible.
Mechanistic accounts in science commonly describe entities and activities organised so that they produce a phenomenon.[5] The constitutive question is selective. A rack screw is inside the server but usually irrelevant to the decision capacity. A power feed is causally necessary for operation but may be an enabling condition rather than a component of the computational mechanism. A training corpus caused the checkpoint to acquire its parameters but is not thereby part of the runtime inference process.
A cause of the episode need not be a component of the capacity. Conversely, a component may be constitutively relevant even when its contribution is distributed and no single episode can be attributed to it alone.
| Role | Question it answers | Typical test | Common overclaim |
|---|---|---|---|
| Indicator | What measurement accompanies Y? | Association, calibration, out-of-sample prediction | “The indicator produces Y.” |
| Trigger or cause | What difference changes Y? | Intervention, natural experiment, identified counterfactual | “Whatever changes Y is part of Y.” |
| Enabling condition | What must remain available for Y to occur? | Removal, degradation, resource constraint | “Necessity proves constitutive role.” |
| Constituent | What organised activity realises Y? | Selective perturbation, organisation change, replacement, restoration | “One successful ablation reveals the whole mechanism.” |
Two clocks of explanation
Causal explanations usually connect differences across time, even when the interval is tiny. A configuration at one moment changes an outcome later. Constitutive explanations are commonly treated as synchronous at the relevant grain: the organised parts and the realised phenomenon occupy the same operating episode. Petri Ylikoski argues that constitution is a synchronous, asymmetric dependence, while developmental explanations can combine causal and constitutive dependencies across time.[8]
The word “synchronous” should not be mistaken for a frozen instant. A handshake protocol, oscillation or recurrent computation unfolds. The appropriate grain may be a 200 ms cycle, a complete transaction or a conversation turn. The rule is more practical: historical production and present realisation must not be collapsed into one relation. Training caused a model state. Runtime computation realises an output. Deployment configuration determines which larger system realises an authorised action.
Intervention is necessary but not self-interpreting
Mechanistic research often uses interventions across levels. Change a component and observe the whole phenomenon. Change the whole operating condition and inspect component activity. Craver’s mutual manipulability proposal tried to capture this practice as a test of constitutive relevance.[6] The appeal is clear: a genuine component should matter to the whole, and the whole’s operating states should be reflected in its components.
The difficulty is equally clear. Interventions designed for causal relations between independent variables behave strangely when one variable is part of the other. Changing the whole without changing its parts may be impossible. Changing a part may affect many things at once. Baumgartner and Gebharter call attention to “fat-handed” interventions, and later work has continued to debate how matched interlevel experiments support constitutive inference.[7][9] There is no settled universal algorithm that turns a perturbation into a constitution verdict.
That does not make constitutive explanation arbitrary. It changes the standard of evidence. Ablation shows that a system can be broken at X; reconstitution tests whether the proposed organisation can make the capacity return. A strong case combines selective perturbation, organisation change, replacement, recovery and boundary comparison. No single positive result carries the full burden.
One incident, three explanations
Return to the opening incident. The trace review finds that 41 of 43 held applications had confidence below 0.45. That is useful operational evidence. It identifies a compact predictor and tells investigators where to look. It still leaves several worlds open: low confidence may trigger the hold, hidden case difficulty may produce both, or selection after automated review may manufacture the pattern.
The lever review then separates three changes. Rebuilding the stale evidence index raises the score from 0.38 to 0.71 for the focal application. Holding the evidence fixed and lowering the policy threshold from 0.45 to 0.35 releases it. Holding both score and threshold fixed while removing the caller’s authority prevents any state change. These interventions identify different causal roles. The index affected the proposal, the threshold affected the decision, and authority affected whether the decision could become an operational effect.
Those results explain why this hold occurred, but they still do not explain what an operational hold is in the deployed system. At the assembly depth, the phenomenon is defined as an authoritative account state that blocks a downstream transaction and can be verified by readback. The relevant operating window begins with evidence retrieval and ends when the system of record confirms the state. Within that window, the model emits a proposal. The policy converts typed fields into a decision. The authority kernel checks identity, purpose, permission and freshness. The state service writes the hold. Readback verifies that the intended effect exists.
Now run replacement tests. Replace the neural scorer with a deterministic rule that emits the same 0.38 for this evidence. The hold is still realised. The neural architecture was causally responsible for the original score, but it was not necessary to the larger capacity once its functional role was preserved. Replace the authority kernel with a logger that always returns “approved” but cannot confer permission. The screen can still display “hold”, yet no authoritative state changes. Behaviour at the interface is preserved while the operational phenomenon disappears.
This gives three bounded conclusions. At trace depth, low confidence was associated with holds in a selected incident set. At lever depth, the stale index, policy threshold and authority checks each made causal differences along the focal trajectory. At assembly depth, the operational hold was realised by an organised chain that included policy, authority, state transition and verification. The model was a proposal-producing component of that chain, not the whole decision-maker.
The example also shows why boundary choice matters. At the model boundary, retrieval is an external input and policy is downstream environment. At the decision-service boundary, retrieval, proposal and policy may all be internal. At the governed-operation boundary, authority, world state and readback enter the constitutive account. None of these boundaries is automatically correct. The right one is the smallest boundary that contains the phenomenon named in the claim.
A decision service reads a local evidence index. Move the same index to a remote service while preserving contents, latency, interface and failure behaviour. The model’s outputs remain unchanged. Is retrieval still part of the decision system?
If the target is the transformer’s token computation, retrieval is an external cause of its input. If the target is the configured account-review service, retrieval may remain a constituent despite crossing a machine boundary. Physical location did not settle the issue. The explanatory target, control relations, operating organisation and chosen system boundary did.
Replace a neural policy module with a lookup table that matches every previously tested input and output. Surface behaviour is preserved. Now intervene on an internal feature that the neural model used. The lookup table has no corresponding feature, so the systems diverge under the new test.
Observed equivalence supported a functional claim over the tested cases. It did not establish shared mechanism. A stronger claim needs an intervention mapping that preserves the relevant counterfactual structure, not only a list of outputs.
Interventions that lie
An intervention is evidence only for the contrast it actually creates. It can mislead when it changes several pathways, targets a proxy, destroys the operating regime or assumes the system boundary it was meant to discover.
The clean scalpel and the sledgehammer
Consider a transformer head whose ablation reduces performance on a reasoning task. The result establishes that the intervention made a difference. It does not immediately establish that the head stores “the reasoning rule”. Zeroing the head may alter activation scale, residual balance, later attention patterns and decoding. The semantic interpretation requires additional tests: targeted activation patching, alternative baselines, path-specific interventions, restoration, task controls and replication across prompts.
Mechanism claims are especially vulnerable to destructive interventions. Pulling a power cable disables every software capacity, yet electricity is not a constituent of each algorithm in the same explanatory sense as its state transitions. Deleting a database index may change the result by causing timeouts rather than by removing the represented evidence. Necessity under damage is weaker than relevance under normal organisation.
Boundary tests are interventions on the explanation
A system boundary is not just a line on an architecture diagram. It controls which dependencies count as inputs, context, components and environment. Move episodic state from a process-local store to a network service. The causal dependency may remain. The constitutive claim may change at one boundary and remain stable at another. That is why “the model remembers” is often an untyped statement. The checkpoint, prompt, external store, world state and evidence ledger have different roles.
A useful boundary test asks whether the proposed phenomenon can still be specified when the candidate component is relocated, replaced or shared. If the same external service simultaneously supports many agents, is it part of each, part of a larger joint system, or infrastructure that causally enables all of them? Evidence will not answer until the target capacity and individuation rule are explicit.
Two agents read and write the same persistent notebook. Each fails a long-horizon task when the notebook is removed. Next, give each a private copy synchronised every minute. Performance remains unchanged on current tests, but conflicts appear when both act within the same minute.
The original ablation showed causal dependence. The synchronisation intervention reveals that the shared notebook also helped constitute a coordination process at the multi-agent boundary. It did not prove that the notebook was part of either model’s individual memory.
When a kill switch proves too much
A team wants to identify the mechanism by which a service compares two account records. They ablate the comparison module and accuracy collapses. That seems informative. They then disconnect the power supply and accuracy also collapses, this time to zero. A crude necessity rule would rank the power supply as the most important comparison component because its removal has the largest effect.
The absurdity reveals the missing contrast. Power is a generic enabling condition for every process on the machine. Its removal does not selectively alter the relation between fields, matching rules and comparison outputs. Replace mains power with a battery while preserving voltage and timing, and the comparison continues unchanged. The replacement test supports the classification of power as an interchangeable resource at the chosen computational boundary, not a phenomenon-specific constituent.
Now disable a checksum daemon. The service again stops producing comparisons because the gateway fails closed when integrity status is absent. The daemon has a real causal effect on observed throughput. It may even be part of the governed operation at a wider security boundary. Yet moving checksum validation to an external gateway while preserving the same signed contract leaves the record-comparison capacity intact. The daemon is therefore not supported as a constituent of the comparison algorithm, although integrity verification may constitute the authorised service transaction.
Finally, patch only the candidate comparison representation while preserving activation scale, timing and the gateway path. Specific match judgements change in the direction predicted by the proposed rule. Restore the original representation and the judgements recover. Replace the module with a different implementation that preserves the same intervention structure, and the capacity survives. This pattern is stronger than destructive necessity because it combines selectivity, prediction, replacement and recovery.
The negative case matters as much as the positive one. Suppose restoration fails even though the candidate representation is returned. Then the original ablation probably altered hidden state or downstream organisation. The correct conclusion is not “the mechanism is mysterious”. It is that the intervention did not isolate the claimed component. Large causal impact cannot compensate for an unspecified contrast.
From internal effect to mechanistic explanation in AI
Causal representation learning asks learned variables to support intervention, transfer and changes of regime rather than merely compress observed regularities.[10] Causal abstraction offers a useful bridge from that demand to mechanistic interpretation. A high-level algorithmic model is faithful to a lower-level network only when interventions correspond across levels and preserve the relevant effects. Recent work has generalised this framework across many interpretability methods, including activation and path patching, causal tracing, concept erasure and sparse feature interventions.[11] The contribution is not that every successful patch discovers “the” mechanism. It is a language for asking whether a proposed high-level story commutes with lower-level interventions.
The mapping itself needs discipline. If an alignment function is allowed to become arbitrarily expressive, it can map a network to an algorithm in ways that cease to be informative. Current results on the non-linear representation dilemma make this break condition explicit.[13] Intervention faithfulness requires constraints on what counts as the same variable across levels. Otherwise, the interpretation can be fitted after the fact.
The exact metaphysics of constitution remains disputed. Some accounts keep constitution non-causal and synchronous. Others analyse it through matched causal relations, causal betweenness or diachronic dynamics. Recent debate continues to challenge whether available experiments provide direct or only indirect evidence of constitutive relevance.[12]
The practical framework below therefore produces graded, typed conclusions. It never outputs “constitution proved”.
A depth-aware practice
The practical mistake is usually made before analysis. A team starts with a dataset or ablation and only later decides what kind of explanation it wanted. Reverse the order. Type the claim first, then design the evidence.
The observation–intervention–constitution assay
The proposed assay is a operational synthesis, not a new metaphysical theory. It treats explanation as three cuts through the same case. Each cut adds fields that the previous one could leave unspecified.
| Depth | Required specification | Positive result permits | Still prohibited |
|---|---|---|---|
| Trace | Measured variables, population, selection, time order, predictive validation | “X is associated with Y in regime R.” | Direction, causal effect, componenthood |
| Lever | Intervention, comparator, estimand, assignment logic, confounders, off-target paths | “Setting X changes Y under conditions C.” | Actual cause in every case, full mechanism, constitution |
| Assembly | Phenomenon, boundary, time grain, organised role, selective perturbation, replacement and recovery | “X is a supported constituent candidate in mechanism M.” | Metaphysical identity, exhaustive mechanism, transfer to another boundary |
At the trace depth, the main adversaries are measurement error, selection and unstable distributions. At the lever depth, they are confounding, interference, ambiguous treatment and poor surgicality. At the assembly depth, they are arbitrary boundaries, historical causes masquerading as current parts, generic enabling conditions and interventions that destroy the organisation they were meant to reveal.
Every depth should emit both a permitted conclusion and a prohibited inference. This one design choice prevents claim migration. It also makes negative results useful. A failed randomisation does not erase an association. A failed reconstitution does not erase a causal effect. It tells you exactly where the explanatory ladder stops.
A constituent claim should strengthen when three intervention families converge: selective perturbation changes the phenomenon as predicted; replacement or relocation preserves the relevant organisation and capacity; reconstitution restores the capacity after controlled disruption.
The proposal is weakened when only destructive ablations work, many generic resources produce the same pattern, the result disappears under a plausible boundary, or restoration succeeds without the alleged component. A rival explanation is that the candidate is a bottleneck for timing, energy or access rather than part of the phenomenon-specific mechanism. The smallest useful implementation is the worksheet below.
type ClaimDepth = "association" | "causal" | "constitutive";
type EvidenceStatus =
| "observed"
| "identified"
| "intervened"
| "triangulated"
| "contested";
interface ExplanatoryClaim {
targetPhenomenon: string;
candidateDependency: string;
depth: ClaimDepth;
populationOrSystem: string;
boundary: string;
timeGrain: string;
intervention?: {
setter: string;
comparator: string;
assignmentLogic: string;
offTargetPaths: string[];
};
organisationEvidence?: {
componentRole: string;
replacementTest: string;
restorationTest: string;
rivalBoundary: string;
};
evidenceStatus: EvidenceStatus;
permittedConclusion: string;
prohibitedInference: string;
}
// Invariants
// 1. causal requires an intervention or an identification argument.
// 2. constitutive requires a declared boundary and organised role.
// 3. every record must state what the evidence cannot establish.
The schema exposes a logic that prose often hides: the same observation cannot silently change its claim depth, and every deeper claim must carry additional fields.
Type the claim before judging the evidence
Complete the fields, then generate a bounded conclusion and a typed JSON record. Use “Load worked example” to inspect the synthetic account-review incident.
Worksheet result
Complete the worksheet and select “Evaluate claim”. The result will state the deepest warranted claim, missing evidence and prohibited inference.
{
"status": "awaiting input"
}
What each result means
A positive association result permits prediction inside the validated regime. It does not justify changing X. A positive causal result permits an action claim for the stated intervention, population and outcome window. It does not prove that X is the only cause, the actual cause of every episode, or a component of the target system. A positive constitutive assessment permits design and research attention to X as part of a specified mechanism. It does not establish metaphysical identity or exhaust the mechanism.
A negative result is diagnostic. If the observation is weak, improve measurement. If causal identification fails, redesign assignment or collect different data. If constitution evidence fails, clarify the boundary, seek selective interventions, attempt replacement and restoration, or retreat to a causal claim. Retreating to the claim the evidence supports is progress, not defeat.
Optional depth: assumptions encoded by the worksheet
The causal score treats randomisation or a stated identification strategy as stronger than natural-group comparison. It penalises high off-target risk and an absent rival generator. This does not calculate statistical significance or identify a causal effect from raw data.
The constitution score requires causal support plus boundary inclusion, an organised role and at least one of replacement or restoration. It gives its strongest result only when both are present and off-target risk is low. This is a practitioner decision rule motivated by mechanistic research practice, not a proof of a disputed metaphysical relation.
The worksheet assumes that the target phenomenon is sufficiently well individuated to support intervention. It should not be applied unchanged to abstract mathematical explanation, logical grounding or cases where intervention itself destroys the phenomenon’s identity.
The decision this changes
The next time a metric, lesion, ablation or patch “explains” a result, do not ask whether the evidence is impressive. Ask which decision it is meant to support.
For an alarm, ranking rule or forecasting system, stable correlation may be enough. For changing policy, treatment or architecture, you need a causal estimand and a defensible intervention or identification strategy. For replacing a module, locating a cognitive boundary, interpreting an internal circuit or claiming that a machine property constitutes a mental capacity, you need an organised mechanism model with boundary, time grain, selective intervention and recovery evidence.
Use correlation to anticipate, causation to change, and constitution to design or identify. Never promote a claim merely because the next word sounds more explanatory.
The practical consequence reaches beyond terminology. It changes incident reviews, model interpretation, scientific experiments and arguments about artificial minds. A report may correlate with an internal state. An intervention may show that the state causally controls the report. Neither alone establishes that the state constitutes experience, selfhood or agency. That further claim needs a declared theory, candidate boundary and measurement bridge.
Explanation advances when the question and the test fit. The archaeological discipline is simple: read the trace, pull the lever, expose the assembly, and record where the evidence stops.
Compact glossary
- Association
- A dependency in a measured distribution. It may support prediction without identifying direction or mechanism.
- Causal effect
- A contrast between outcomes under specified alternative assignments or interventions.
- Constitution
- A part–whole or organisation–phenomenon dependence concerning what realises a capacity, state or process.
- Estimand
- The exact causal quantity sought, including intervention, comparator, population and outcome window.
- Fat-handed intervention
- An intervention that changes several relevant pathways at once, weakening interpretation.
- Identification
- An argument that a causal estimand can be recovered from available data under stated assumptions.
- Mechanism
- Entities or components and activities organised so that they produce or realise a phenomenon.
- Reconstitution
- Restoring or rebuilding a proposed organisation to test whether the target capacity returns.
- System boundary
- The declared division between the candidate system and its environment for a specified explanatory target.
- Time grain
- The temporal window at which the phenomenon and its realising organisation are individuated.
Primary and authoritative references
Open the source register and extended notes
- Pearl, J. (1995). “Causal diagrams for empirical research.” Biometrika, 82(4), 669–710. DOI and publisher record.
- Rubin, D. B. (1974). “Estimating causal effects of treatments in randomized and nonrandomized studies.” Journal of Educational Psychology, 66(5), 688–701. DOI.
- Hernán, M. A., & Robins, J. M. (2020, living edition). Causal Inference: What If. Official open book and materials.
- Woodward, J. (2002). “What is a mechanism? A counterfactual account.” Philosophy of Science, 69(S3), S366–S377. DOI.
- Machamer, P., Darden, L., & Craver, C. F. (2000). “Thinking about mechanisms.” Philosophy of Science, 67(1), 1–25. DOI.
- Craver, C. F. (2007). “Constitutive explanatory relevance.” Journal of Philosophical Research, 32, 3–20. DOI.
- Baumgartner, M., & Gebharter, A. (2016). “Constitutive relevance, mutual manipulability, and fat-handedness.” British Journal for the Philosophy of Science, 67(3), 731–756. DOI.
- Ylikoski, P. (2013). “Causal and constitutive explanation compared.” Erkenntnis, 78(2), 277–297. DOI.
- Craver, C. F., Glennan, S., & Povich, M. (2021). “Constitutive relevance & mutual manipulability revisited.” Synthese, 199, 8807–8828. DOI.
- Schölkopf, B., Locatello, F., Bauer, S., et al. (2021). “Toward causal representation learning.” Proceedings of the IEEE, 109(5), 612–634. DOI.
- Geiger, A., Ibeling, D., Zur, A., et al. (2025). “Causal abstraction: a theoretical foundation for mechanistic interpretability.” Journal of Machine Learning Research, 26(83), 1–64. Official article and PDF.
- Kistler, M. (2025). “Constitution and causation in mechanisms.” Análisis Filosófico, 45(Special issue), 571–602. DOI.
- Sutter, D., Minder, J., Hofmann, T., & Pimentel, T. (2025). “The non-linear representation dilemma: is causal abstraction enough for mechanistic interpretability?” Primary preprint.