A tone nobody agrees was heard
A participant says she heard nothing. Her forced-choice answer identifies the tone correctly. Her pupil expands. An interviewer later elicits a fleeting “pressure” before the button press. Which record is evidence of what?
Imagine a laboratory testing near-threshold hearing. On one trial, a soft tone plays through headphones. Asha selects “left ear” correctly but reports no auditory experience. Her confidence is low. A pupil camera registers dilation. A late scalp potential appears. Twenty minutes later, a skilled interviewer asks her to return to the instant before the choice. Asha describes a brief directional pull that did not feel like a sound.
Five teams can now tell five confident stories. The phenomenologist says the “pull” was the experience. The psychophysicist says above-chance discrimination shows perception without awareness. The metacognition researcher says low confidence reveals poor access to a correct first-order judgement. The neuroscientist points to the potential. The mechanistic modeller fits an evidence accumulator whose state crossed a decision threshold.
Each story might be useful. None follows from the raw record alone. The report, interview, button press, pupil trace and model parameter were produced through different access routes. They also answer different questions. The central mistake is to treat convergence among outputs as identity among what those outputs measure.
Part IThe evidence changes as it travels
First-person evidence begins from a privileged relation: only Asha can directly undergo whatever occurred for her. That privilege does not make every later sentence complete, causally accurate or immune to framing. A report is an action performed after attention, memory, categorisation and language have already shaped what can be said. Classic work on verbal reports found poor access to many higher-order cognitive causes, while a different tradition showed that carefully constrained concurrent reports can be informative about attended information.[2][3] The useful conclusion is not that introspection is either infallible or worthless. It is that the report-generation process must itself be modelled.
Three claims hide inside one report
Take the sentence “I heard a faint tone”. It can support at least three different claims. The first is historical: Asha now reports that she heard a tone. The second is experiential: a tone-like episode occurred for her at the relevant moment. The third is causal: a specific sensory process produced that episode. The first claim follows directly from the recorded utterance. The second requires a bridge from present report to prior experience. The third requires a further bridge from experience to mechanism. A single sentence can therefore be direct evidence for one claim, indirect evidence for another and almost silent about a third.
This separation prevents two symmetrical errors. The sceptical error treats every transformation as corruption and concludes that reports reveal nothing. The credulous error treats privileged access as if it survived attention, retention and expression without loss. A better model asks where information can disappear, where new structure can enter and which changes should leave the report invariant. The report remains indispensable because it may be the only channel that identifies a phenomenal distinction. It remains fallible because identification is not the same as complete recovery.
The same discipline applies to negative reports. “I experienced nothing” may mean no relevant content was available, no stable category could be formed, the scale offered no suitable option, the memory decayed, or the participant adopted a conservative criterion. Absence of a report is an observation about a report channel before it is evidence about absence in experience. A study earns the stronger inference only by excluding plausible failures in access, retention and expression.
Interpersonal evidence is often hidden inside the phrase “self-report”. Yet a response to a checkbox, a clinical interview and a micro-phenomenological elicitation are not the same operation. A trained interviewer can help a participant return to a specific episode, distinguish sequence from interpretation and notice dimensions that an unaided summary would omit.[7] The exchange is productive precisely because another person shapes attention and language. That means the interaction is part of the measurement apparatus.
Sometimes the interaction is more than an apparatus. Embarrassment, mutual gaze, trust repair and conversational alignment are partly constituted by reciprocal exchange. Research that replaces interaction with passive observation can therefore remove part of its target. Second-person neuroscience developed from this point: live reciprocal interaction may recruit processes not captured by watching social stimuli from outside.[14]
Probe, perturbation or constituent?
An interpersonal method can play three causal roles. As a probe, it seeks information about an episode without materially changing the relevant feature. As a perturbation, it changes attention, memory, categorisation or affect and thereby changes the target. As a constituent, the interaction is part of the phenomenon being studied. The role may change within one interview. The evidence manifest must declare which role is assumed at each step.
Third-person evidence is also plural. Behaviour records an organism or system acting under a task, incentives, abilities and policies. Physiology records a signal after transduction, filtering, localisation and preprocessing. A mechanism claim proposes a causal organisation that should survive intervention. These are not three names for objectivity. They are three different observation functions.
| Evidence class | Direct record | Common overreach | Strong next test |
|---|---|---|---|
| First-person report | What can be accessed, retained, discriminated and expressed under the reporting conditions | Treating the report as a complete copy of experience or a correct account of hidden causes | Vary timing, vocabulary and report demands while holding the episode as stable as possible |
| Interpersonal elicitation | A jointly produced description and the dynamics of producing it | Treating interviewer-assisted detail as either pure discovery or mere suggestion | Preserve prompt lineage, compare interviewers and test invariants across neutral re-elicitation |
| Behaviour | Action under a task, policy, incentive and motor channel | Equating correct action with awareness, intention or felt confidence | Change payoff, response mapping or motor route without changing the stimulus |
| Physiology | Instrument output linked to biological dynamics through a measurement model | Reading a cognitive or experiential label directly from a signal | Test selectivity, temporal order and perturbation response across tasks |
| Mechanism | A model of causal parts, relations and transitions | Treating fit, decoding or correlation as evidence that the proposed parts do the work | Ablate, intervene, restore and compare a rival generator |
The rewarded denial
Repeat Asha’s tone trial, but pay a large bonus whenever she answers “unheard”. Her experience may remain unchanged while her report changes. Now reverse the reward and make the button mapping awkward for her injured hand. Behaviour changes again. If the pupil trace remains similar, the three lenses diverge.
The divergence is not a methodological failure. It identifies two hidden variables, report incentive and motor cost. A theory that equated report, behaviour and experience would have no place to put them. A typed theory predicts the divergence before it occurs.
Part IIEvery lens has an observation function
A useful measurement model begins with a modest claim. Suppose there is a candidate episode, X, such as Asha’s momentary auditory experience. Method k does not deliver X directly. It produces an observation Yk through a transformation that includes access conditions, context and noise.
For a report, the transformation includes what was attended, how long the delay lasted, what distinctions the participant can make and which words the protocol offers. For an interview, it includes the question sequence, relationship, expectations and mutual corrections. For behaviour, it includes the decision policy, payoff and response channel. For physiology, it includes the biological process, sensor geometry and analysis pipeline. The noise term is not the only problem. A perfectly repeatable transformation can still measure the wrong thing.
What a bridge must carry
A bridge is not the assertion that two measures correlate. It states why an observation should change when the target changes, why relevant alternatives should not produce the same change and where the relation is expected to fail. For Asha’s pupil trace, a usable bridge might say that auditory evidence increases pupil-linked arousal within a specified window after controlling luminance, surprise, effort and report preparation. That bridge is narrow, conditional and testable. “The pupil reveals awareness” is not.
Every bridge should name four elements: the target property, the transformation that exposes it, the nuisance variables that can imitate it and the intervention that separates the preferred path from its rivals. A measurement bridge is strong when it survives the variation it claims should not matter and breaks under the variation it claims should matter. Reliability alone cannot supply this. A biased scale can return the same number every time. Construct validity grows from a network of discriminating relations, not from repeatability in isolation.
Optional technical depth: observation functions and identification
The equation above is deliberately permissive. It does not assume that the candidate episode is a scalar, that noise is additive, or that every method observes the same latent variable. In a fuller model, each channel may depend on several hidden states and may feed back into later states. An interview at time two can alter memory at time three. A report instruction can alter attention before the stimulus. A physiological analysis can select a window after seeing behavioural results.
Identification requires more than fitting the observations. Rival parameterisations or rival causal graphs may generate the same joint distribution. Method variation and targeted intervention supply additional constraints. The practical question is not whether one latent model explains the data, but whether the configured study makes the preferred bridge distinguishable from credible alternatives.
This yields a compact evidence unit: observation + bridge + claim + rival generator. “Pupil dilation occurred” is an observation. “Pupil dilation indicates conscious hearing” adds a bridge and a claim. Arousal, surprise and effort are rival generators. Without them, the sentence looks precise while hiding the step that matters.
One faint tone, five records
Suppose a synthetic trial produces these values: report = “unheard”; confidence = 0.25; left-right discrimination = correct; pupil change = +0.18 standard units; late EEG component = +2.1 μV. The temptation is to average them into a consciousness score. That operation destroys their meaning.
The correct discrimination is evidence that information relevant to ear location influenced action. It is not yet evidence that the tone was experienced. Low confidence records a metacognitive judgement, not the absence of first-order evidence. Pupil and EEG values are instrument outputs whose selectivity must be established. The elicited “directional pull” is evidence about a later-accessible experiential distinction under that interview protocol. The mechanistic model must predict how these records change when signal strength, payoff and reporting requirements are independently manipulated.
Timing is therefore part of the evidence type. An immediate confidence judgement, a retrospective explanation and a next-day narrative should not occupy one column called “subjective data”. They have different access paths. Micro-phenomenological analysis makes this explicit by reconstructing diachronic and synchronic structures from concrete episodes and tracking the analysis process.[8] Descriptive Experience Sampling similarly aims to anchor reports to sampled moments, and interobserver reliability can be evaluated separately from whether the description is true of the experience.[6]
The helpful interviewer
Two interviewers explore the same trial. One repeatedly offers visual language: “Was it a flash, a shape, a brightness?” The other asks only for temporal sequence and allows uncertainty. Asha later reports a “small flash” to the first and a “directional pressure” to the second.
Three explanations remain open. The first interviewer may have supplied a label for a real but previously inexpressible aspect. The prompt may have reconstructed the memory. Or the two descriptions may pick out different features. The study cannot decide by preferring the richer transcript. It needs prompt lineage, neutral re-elicitation, interviewer blinding and a prediction about which features should remain invariant.
Confidence is not a second copy of accuracy
Asha’s confidence may track whether her decisions are correct, but confidence level and metacognitive sensitivity are distinct. A person can be generally overconfident yet discriminate relatively well between correct and incorrect trials. Measures such as meta-d′ were developed to separate metacognitive sensitivity from first-order performance and response bias.[11] This matters because a low confidence rating cannot simply be counted as a negative version of a correct button press.
The confidence trap
Consider two participants with 75 per cent tone-location accuracy. Asha uses the full confidence scale and is usually more confident on correct trials. Dev gives “high confidence” on almost every trial. Their average confidence may be identical, but Asha’s confidence contains more trial-level information about correctness. If a study equates mean confidence with awareness, it misses this distinction. If it treats accuracy as awareness, it misses the possibility of correct performance with weak access. The two measures constrain one another only after their separate generative processes are represented.
Third-person does not mean interpretation-free
Brain and body measurements often feel less negotiable because a machine records them. Yet an EEG component is created through electrode placement, reference choice, filtering, epoch selection, artefact rejection and statistical contrast. A region’s activation is not a deductive label for a cognitive process. Reverse inference depends on how selectively that activity occurs across alternative processes and tasks.[12] Recent hyperscanning work has shown that plausible analysis choices can substantially alter estimates of inter-brain synchrony, a vivid reminder that interpersonal physiology is not raw relational truth.[19]
No-report paradigms respond to one important confound. If asking for a report recruits decision, working memory and motor preparation, removing the report may help separate those consequences from earlier activity. But a no-report design does not observe experience without a bridge. It substitutes another indicator, such as involuntary eye movement, pupil response or stimulus intensity, and inherits the assumptions of that indicator.[13] A current review of human intracranial research makes the same boundary explicit: no-report approaches reduce some post-perceptual confounds, but they still rely on indirect classification, and null effects remain difficult to interpret.[16]
The silent pianist
A pianist experiences a vivid inner melody but cannot move or speak because of temporary paralysis. Behavioural evidence disappears while first-person experience, if later reported, may remain rich. Now consider an automated player piano that reproduces the same keystrokes from a stored file. Behaviour matches, but the causal organisation differs.
The case does not prove that the pianist is conscious and the machine is not. It proves something narrower and crucial: output matching under one interface cannot identify experience or mechanism. The missing channel creates uncertainty, not a negative verdict.
Part IIIDisagreement is where the causal work begins
Researchers often use “triangulation” to mean that several measures point in the same direction. That is too weak. Measures can agree because they share a confound, a vocabulary, a preprocessing choice or a demand characteristic. Three thermometers placed in the same patch of sunlight do not independently establish the room temperature.
Reciprocal constraint begins when one channel makes a risky prediction about another under a controlled change. A first-person distinction should predict a behavioural or physiological contrast beyond generic arousal. A mechanistic intervention should predict a specific experiential change, not merely reduced performance. An elicitation category should remain recognisable when wording, interviewer or timing changes, within stated limits. A physiological marker should distinguish the target from relevant alternatives across tasks.
Disagreement has more than one meaning
When two channels diverge, the first task is to type the disagreement. A translation disagreement occurs when two vocabularies divide the same episode differently. A temporal disagreement occurs when methods sample different moments. A policy disagreement occurs when report or action is changed by incentives, confidence or motor cost. A mechanism disagreement occurs when matched observations arise from different causal organisations. These cases require different repairs. More participants cannot repair a mistranslated construct, and a better sensor cannot repair a criterion shift.
Disagreement should lower confidence when a claimed bridge predicts agreement under the configured conditions and no recorded nuisance explains the failure. It should increase information when rival explanations predict different patterns and the observed divergence selects among them. Parallax becomes evidence only when the expected direction of displacement was stated before the result was known. Otherwise, every mismatch can be redescribed after the fact as methodological richness.
This is why the unit of progress is not “three methods agreed”. It is a prediction delta: under intervention I, report should change by this kind of amount, behaviour should remain stable within this range, and the proposed physiological marker should shift in this time window. Cross-constraint is directional and conditional, not a vote among methods. One method may constrain the interpretation of another without being measured on the same scale or assigned equal evidential weight.
Change the target and the lens separately
There are four basic moves. An invariance test changes the measurement method while trying to hold the target stable. An intervention test changes the target or proposed mechanism while holding the method stable. A dissociation test changes a suspected nuisance while holding the target stable. A recovery test removes the nuisance and asks whether the predicted relation returns. Together, they turn disagreement into diagnostic evidence.
False convergence from common demand
Figure 5 comes from a simple synthetic model of 160 trials. Report, behaviour and pupil response each receive a large contribution from task demand and only a small contribution from the intended target. Their pairwise correlations range from 0.84 to 0.86. Yet their correlations with the target range from -0.04 to 0.10.
A pooled score would look stable and multimodal. It would be stably wrong. The revealing intervention is to vary demand while holding the target-generating stimulus constant. If all three measures move together, the shared movement is evidence about demand sensitivity. It does not become evidence about the target by occurring in three columns.
Four forms of legitimate cross-constraint
- Phenomenology to experiment. A recurring experiential distinction generates a new contrast, temporal marker or task condition. Lutz et al. used participant-reported cognitive contexts to form phenomenological clusters, then examined whether distinct clusters were associated with different large-scale EEG synchrony patterns.[10] The correlation did not turn reports into EEG or EEG into experience. It made each analysis conditional on the other.
- Experiment to elicitation. A behavioural or physiological transition identifies a narrow time window for later interview. The interviewer asks about that episode without disclosing the physiological classification. The aim is not confirmation but a prediction: trials from one hidden class should contain a particular experiential sequence more often than matched alternatives.
- Mechanism to lived consequence. A perturbation changes a proposed causal component and predicts a specific shift in timing, structure or accessibility of experience. A generic reduction in accuracy is insufficient because fatigue, distraction or motor impairment could produce it.
- Method to method. The same claimed feature survives changes in interviewer, scale, response mapping or sensor pipeline that should alter method artefacts. This is convergent validity with an explicit model of method effects, extending the logic of multitrait-multimethod analysis.[4]
The perfect mimic
Build two systems. One integrates sensory evidence over time and uses an internal confidence variable to govern report and action. The other reads a lookup table created from the first system’s recorded outputs. On the original test set, their reports, confidence ratings and actions match exactly.
A behavioural comparison cannot separate them. A self-description prompt may not separate them either, because both can emit the same text. A state intervention can. Swap or perturb the accumulator state in the first system and its later outputs change in a structured way. The lookup system has no corresponding state. Mechanism evidence enters when an intervention distinguishes matched generators. Even then, the result establishes a causal organisation, not experience, unless an additional theory supplies that bridge.
A protocol for reciprocal constraint
- Type the claim. Is it about reported content, experiential structure, task information, biological correlation, causal mechanism or metaphysical status? One sentence should contain only one type.
- Fix the candidate boundary and episode. Specify the person, dyad or system, the temporal window and what counts as the target event. Do not let “mind” float between organism, interaction and instrument.
- Write each observation function. Record access, delay, prompt, incentive, sensor, preprocessing and model-fitting decisions that transform the target into data.
- Draw the dependence graph. Mark shared prompts, shared trials, shared coders, shared preprocessing and common causes. Multiple columns are not independent when they inherit the same source.
- Pre-register prediction deltas. State how each measure should change under target, nuisance and mechanism interventions. “All measures will increase” is rarely discriminating enough.
- Build a rival generator. Specify a process that can reproduce the headline observation without the preferred explanation. Prefer an implemented mimic or an explicit causal model.
- Type the conclusion. State what a positive result permits, what a negative result weakens and which inference remains prohibited.
Part IVBuild the manifest before collecting the data
The multi-method evidence manifest below turns the protocol into a field instrument. It is deliberately not a score. A study can have excellent first-person detail and weak causal discrimination, or strong mechanism evidence and no access to experience. Compressing those states into 73 out of 100 would manufacture an ordering that the evidence does not support.
The manifest instead returns separate statuses for identification, channel coverage, bridge clarity, dependence, discrimination and conclusion boundaries. It can be used during study design, preregistration, model evaluation, clinical protocol review or research governance.
The manifest is a contract, not a data warehouse
The manifest does not need to contain every transcript, waveform or model checkpoint. It records the claims that make those materials usable as evidence: which episode they concern, how they were produced, which transformations are authorised, which dependencies they share and what conclusion may be drawn. The raw artefacts can remain in their native systems. The manifest governs inferential movement between records; it does not erase their provenance by copying them into one schema.
This distinction matters in collaborative research. A phenomenology team may curate experiential categories, a physiology team may define preprocessing and a modelling team may implement rival generators. Without a shared manifest, each team can inherit the others’ labels as if they were ground truth. With one, every hand-off carries a typed claim, its bridge and its unresolved alternatives. A conclusion is ready for review only when another team can reconstruct why each observation bears on it and where that inference must stop.
Executable lab: multi-method evidence manifest
Complete the fields, validate the design, and export a reusable JSON manifest. The validator returns independent statuses rather than a synthetic score.
What the artefact tests
The validator tests whether the study has an identified target, more than one evidential route, an explicit bridge, a declared dependence structure, a rival-generator intervention and bounded conclusions. It encodes one central assumption: methodological strength comes from explicit relations and discriminations, not from the number of measures.
A positive validation result permits the team to say that the design is specified well enough for review and preregistration. It does not validate the empirical claim. A negative result identifies missing design logic. It does not show that the target is absent. The tool cannot establish consciousness, metaphysical status or formal equivalence between an experience and a mechanism.
const checks = {
identified: Boolean(claim && target && boundary),
multiChannel: channels.length >= 2,
bridged: Boolean(bridge),
dependenceMapped: Boolean(dependence),
discriminating: Boolean(rival && intervention),
bounded: Boolean(permitted && prohibited)
};
// No total score is calculated.
// Each false value remains visible as a distinct design debt.
Optional technical depth: minimum manifest invariants
A machine-readable implementation should preserve at least five invariants. Observations retain their native channel type. Every inferential edge names its bridge. Shared causes are represented rather than assumed away. Positive and negative outcomes have separate permitted conclusions. Missing channels remain explicit rather than being converted to zero.
The validator intentionally checks specification completeness, not scientific truth. A team can satisfy every field with weak content. Human and adversarial review must still challenge the bridge, rival generator and intervention. The exported JSON is therefore a review object and versioned research receipt, not a certificate.
A preregistered cross-over study
A small study can use four conditions: weak or strong tones crossed with report or no-report blocks. Immediate report and confidence are collected only in report blocks. Forced-choice location is collected after a delayed cue in all blocks to reduce motor preparation during the critical window. Pupil and EEG are recorded continuously. A separate interviewer, blind to trial class, elicits a sample of specific episodes using neutral temporal prompts.
The preferred mechanism proposes an early sensory-evidence process, followed by metacognitive access and report preparation. The rival says the headline EEG effect is mostly task relevance and motor preparation. The preregistered prediction is not simply that “seen” trials have larger signals. Stronger tones should affect early sensory measures in both block types. Report demand should have a larger effect on late activity and response latency. If an elicited “directional pull” is meaningful, it should recur across interviewers and predict a defined class of trials beyond confidence and demand.
A positive pattern supports a conditional causal decomposition of sensory evidence, access and report consequences. A null early effect weakens that decomposition or exposes insufficient measurement sensitivity. A late-only effect supports the report-consequence rival. None of these outcomes proves or disproves experience itself.
Where a manifest can still fail
A manifest can be complete and the study can still be biased. Teams can write vague bridges, choose easy rivals or declare an intervention that changes several variables at once. Open methods, preserved raw-to-analysis provenance, blinding and independent replication remain necessary. Construct validity is not a property granted by a form. It is an evolving argument connecting a measure to a network of predictions and alternatives.[5]
Reciprocal constraint does not settle ontology
A physicalist, biological naturalist, functionalist, idealist or dual-aspect theory can all use this protocol. They will disagree about what the candidate episode ultimately is and which mechanisms could constitute it. The protocol does not average those positions into neutrality. It forces each position to state its bridge and a result that would weaken it.
First-person evidence is not automatically subordinate because physiology can be measured, and it is not automatic proof of a consciousness-primary metaphysics because experience is directly lived. Third-person mechanism evidence can explain control, report or integration without exhausting subjectivity. The unresolved remainder must stay visible rather than being smuggled into the labels.
Difficult cases
For infants, non-speaking animals, patients without reliable motor output and many artificial systems, a first-person report may be unavailable. The absence of that channel is missing evidence, not evidence of absence. Researchers should show the missingness and narrow the claim to behaviour, physiology or mechanism. Proxy measures may be necessary, but their transfer bridge must be explicit.
For trained contemplatives or expert performers, practice may improve discrimination and vocabulary. It may also alter the phenomenon, attention strategy or demand sensitivity. Expertise should therefore be modelled as both measurement capacity and intervention. For machine-generated self-descriptions, the safe default is to classify them as system outputs with known prompt and training dependencies. Calling them first-person evidence requires a declared theory of the candidate subject, its boundary and why the output has the relevant epistemic relation.
For genuinely interpersonal phenomena, separate individual and dyadic variables. Synchrony can arise from common stimulus timing, shared movement or preprocessing, not only mutual coupling. Conversely, reducing the dyad to two individual brains can miss reciprocal control. The correct unit may be a coupled loop, but that is a hypothesis to test, not a label conferred by simultaneous recording.
The decision this changes
When a claim about mind is uncertain, the usual response is to search for the best measure: the most sensitive questionnaire, the cleanest neural marker, the most natural interview or the most accurate behavioural task. That framing is wrong. No measure escapes its access route.
The better decision is to choose the smallest set of methods with meaningfully different failure modes, then design a case in which they should diverge if a bridge assumption is wrong. Preserve the divergence. Do not erase it with a composite score. Ask which causal change explains the pattern and whether the preferred mechanism predicts recovery.
This changes resource allocation. A fifth correlated measure is often less valuable than one intervention that separates two live explanations. A longer interview is less valuable than a blinded re-elicitation if prompting is the main risk. A higher-resolution brain image is less valuable than a task manipulation if report preparation is the confound. Buy discriminating variation before buying additional agreement.
Decision rule
Before accepting multi-method convergence, specify the intervention that would make the channels diverge if one bridge assumption were false. If no such intervention can be stated, the study has collected several descriptions, not reciprocal constraint.
First-person, interpersonal and third-person evidence become scientifically powerful together because they are not interchangeable. Their differences create parallax. A well-designed study converts that parallax into tests of access, method, nuisance and mechanism. The outcome is not a view from nowhere. It is a disciplined map of what each view can and cannot support.
Glossary
- Bridge assumption
- The claim that explains why an observation bears on a target construct or event.
- Candidate episode
- The bounded event, process or state that the study seeks to understand.
- First-person evidence
- Data produced from a subject’s access to and expression of their own experience under specified conditions.
- Interpersonal evidence
- Data generated through reciprocal interaction, including elicitation, joint attention and co-regulation.
- Observation function
- The transformation from a candidate episode to a recorded datum, including access, context and measurement operations.
- Parallax
- Systematic displacement among observations created by different vantage points and transformations.
- Rival generator
- An alternative process capable of producing the same headline observation without the preferred explanation.
- Third-person evidence
- Externally recorded behaviour, physiology or mechanism tests, each with a distinct measurement model.
References
- Varela, F. J. (1996). Neurophenomenology: A methodological remedy for the hard problem. Journal of Consciousness Studies, 3(4), 330-349.
- Nisbett, R. E., & Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231-259.
- Ericsson, K. A., & Simon, H. A. (1980). Verbal reports as data. Psychological Review, 87(3), 215-251.
- Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81-105.
- Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281-302.
- Hurlburt, R. T., & Heavey, C. L. (2002). Interobserver reliability of Descriptive Experience Sampling. Cognitive Therapy and Research, 26, 135-142.
- Petitmengin, C. (2006). Describing one’s subjective experience in the second person: An interview method for the science of consciousness. Phenomenology and the Cognitive Sciences, 5, 229-269.
- Petitmengin, C., Remillieux, A., & Valenzuela-Moguillansky, C. (2019). Discovering the structures of lived experience: Towards a micro-phenomenological analysis method. Phenomenology and the Cognitive Sciences, 18, 691-730.
- Høffding, S., & Martiny, K. (2016). Framing a phenomenological interview: What, why and how. Phenomenology and the Cognitive Sciences, 15, 539-564.
- Lutz, A., Lachaux, J.-P., Martinerie, J., & Varela, F. J. (2002). Guiding the study of brain dynamics by using first-person data. Proceedings of the National Academy of Sciences, 99(3), 1586-1591.
- Fleming, S. M., & Lau, H. C. (2014). How to measure metacognition. Frontiers in Human Neuroscience, 8, 443.
- Poldrack, R. A. (2006). Can cognitive processes be inferred from neuroimaging data? Trends in Cognitive Sciences, 10(2), 59-63.
- Tsuchiya, N., Wilke, M., Frässle, S., & Lamme, V. A. F. (2015). No-report paradigms: Extracting the true neural correlates of consciousness. Trends in Cognitive Sciences, 19(12), 757-770.
- Redcay, E., & Schilbach, L. (2019). Using second-person neuroscience to elucidate the mechanisms of social interaction. Nature Reviews Neuroscience, 20, 495-505.
- Aru, J., Bachmann, T., Singer, W., & Melloni, L. (2012). Distilling the neural correlates of consciousness. Neuroscience & Biobehavioral Reviews, 36(2), 737-746.
- Stockart, F., Robin, A., Blumenfeld, H., Brázdil, M., Kahane, P., Mudrik, L., Thum, J., Pereira, M., & Faivre, N. (2026). Neural correlates of perceptual consciousness from within: A narrative review of human intracranial research. eLife, 15, RP109604.
- Krakauer, J. W., Ghazanfar, A. A., Gomez-Marin, A., MacIver, M. A., & Poeppel, D. (2017). Neuroscience needs behavior: Correcting a reductionist bias. Neuron, 93(3), 480-490.
- Munafò, M. R., Nosek, B. A., Bishop, D. V. M., et al. (2017). A manifesto for reproducible science. Nature Human Behaviour, 1, 0021.
- Zimmermann, M., Schultz-Nielsen, K., Dumas, G., & Konvalinka, I. (2024). Arbitrary methodological decisions skew inter-brain synchronization estimates in hyperscanning-EEG studies. Imaging Neuroscience, 2.