Pupil and Neural Dynamics Reveal Belief-Dependent Decision Making Under Ambiguity

  1. Department of Biomedical Engineering, Columbia University, New York, United States
  2. Biobehavioral Health, The Pennsylvania State University, University Park, United States
  3. Biomedical Engineering, The Pennsylvania State University, University Park, United States
  4. Department of Electrical Engineering, Columbia University, New York, United States
  5. Department of Radiology, Columbia University, New York, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Angela Langdon
    National Institute of Mental Health, Bethesda, United States of America
  • Senior Editor
    Tirin Moore
    Stanford University, Howard Hughes Medical Institute, Stanford, United States of America

Reviewer #1 (Public review):

Summary:

The manuscript investigates value-based decision-making under risk and ambiguity using a combination of behavioral, pupillometric, and EEG data. Participants are stratified into three "decision styles" (ideal, aggressive, conservative) based on how their choices under known risk align with expected-value optimality. The central claim is that ambiguity aversion is not a uniform bias but reflects heterogeneous internal belief models, and that physiology tracks subjective belief rather than objective task structure. While this is an interesting conceptual question, the evidence is underwhelming given that differences between groups are not tested statistically (but just described), there are clear problems with how the computational models are implemented, and there are serious issues with sampling of participants.

Strengths:

The multimodal design (behavior, pupillometry, EEG) and the attempt to link a latent belief parameter to physiological signatures address a question of clear interest.

Weaknesses:

Framing and motivation

(1) The framing conflates two questions that appear distinct. The motivation centers on "ambiguity aversion," but the study's actual aim - how individuals internally represent ambiguous outcomes - seems like a different question. The relationship between these two framings needs to be made explicit, because as written the motivating phenomenon and the studied phenomenon are not obviously the same thing.

(2) Several of the contrasts the paper sets up against prior literature read as strawmen. The claim that ambiguity aversion is treated as "a single bias or fixed trait that applies uniformly" is presented as the view being overturned, but it is not clear this is a position the field actually holds - it reads as a strawman. Relatedly, the central objective-versus-subjective valuation distinction that the results are built around also reads as a strawman dichotomy rather than a genuine competing account.

(3) The motivation for the physiological measures is overly broad. The statement linking EEG to control, attention, valuation, uncertainty, conflict, effort, and engagement is so general as to be uninformative - EEG signals have been linked to essentially everything, so this does not constrain the hypotheses or predictions. A more specific, falsifiable rationale is needed.
Design, sample, and grouping.

(4) The inclusion of the collaborative spacecraft/Apollo task is unclear. It is not explained why this task is included, and its role relative to the core ambiguity question needs justification (this also bears on the leadership analyses; see below).

(5) The participant numbers do not add up and must be reconciled. The text reports 57 participants, yet the analyses describe three groups of roughly 32 + 32 + 31. The relationship between participants, sessions, and group n's needs to be stated clearly and consistently, because at present the sample description is internally contradictory.

(6) The rationale for categorizing participants into three discrete groups is not established, and the approach is statistically questionable. Decision tendency appears to be a continuous variable; dichotomizing/trichotomizing a continuous measure is generally discouraged and can manufacture or distort group differences. The authors should justify why discrete groups are needed at all, and ideally show whether there are genuine group differences (e.g., evidence of discontinuity/clustering) rather than an arbitrary split of a continuum.

Statistics

(7) Key claims about how ambiguity affects groups differently are made without the appropriate test. To support a claim that the effect of ambiguity differs across groups, the interaction (group × ambiguity) must be shown - group-wise effects reported separately are not sufficient. This is really a key limitation of the current work.

(8) The methods mentioned that some participants performed multiple sessions, but their data were treated as if coming from separate participants. This is incorrect for several reasons, particularly given the focus on individual differences.

Belief parameter and terminology

(9) The term "ideal" is not justified. It is unclear why this group is labelled "ideal" - are they Bayes-optimal, or optimal in some defined sense? If the label implies normativity, that needs to be demonstrated; otherwise it should be renamed.

Drift-diffusion modelling

(10) The boundary parameter is fixed (a detail which is hidden in the methods), but this is highly problematic. By enforcing the same boundary value for all participants, the model is forced to capture any variation as drift rate effects. As such, all conclusions about drift rate are not interpretable as they might reflect boundary effects in disguise.

(11) The DDMs are fit separately per group of participants, which again precludes testing interactions. As with the behavioral analyses, fitting separate models means group differences cannot be properly compared within a single statistical framework, and interactions cannot be assessed. The paper does mention some comparison between groups, but comparing DDM parameter estimates across separately fit models is not valid.

(12) Overall, the DDM is very complex, and the manuscript does not yet provide enough validation to make the model trustworthy. Given the number of trial-wise covariates entering the drift rate and the per-participant fitting, stronger evidence that the model is identifiable and that its parameters are recoverable/reliable is needed before the conclusions drawn from it can be accepted.

Methods - EEG and analysis details

(13) The high-pass filter setting appears very aggressive. The authors should confirm whether this risks removing genuine low-frequency signal of interest, particularly given that delta-band effects are later interpreted.

(14) There is an apparent inconsistency in the epoching/time-locking. The time-frequency analysis appears to be computed on choice-locked data, yet elsewhere the epochs are described as stimulus-locked. This needs to be clarified and made consistent, as it affects interpretation of the pre- versus post-decision EEG clusters.

(15) The mixed-effects modelling appears to omit random slopes. The authors should justify the random-effects structure (e.g., why only random intercepts), as this affects the validity of the inference.

Reviewer #2 (Public review):

Summary:

The manuscript by Qin and colleagues entitled "Pupil and Neural Dynamics Reveal Belief-Dependent Decision Making Under Ambiguity" examines decision-making under risk and ambiguity using pupillometry and EEG. The study employs a lottery choice task with three levels of ambiguity (zero, low, high). Participants were classified into three groups based on their choice behavior in a condition with risk and no ambiguity: ideal (choosing in line with objective expected values), aggressive (preference for investments), and conservative (preference against investments). The authors then compared behavior, pupil, and EEG results across these groups. The study concludes that individual beliefs about ambiguity are reflected in different behavioral strategies and neural correlates.

Strengths:

The combination of behavior, computational modeling, pupillometry, and EEG.

Weaknesses:

(1) It is unclear whether group definition is theoretically justified.

One general concern is that the strategy to form three distinct groups is not clearly motivated. The authors created the three groups, "aggressive", "ideal", and "conservative", based on the zero-ambiguity trials. However, as the authors state: "Ambiguity differs fundamentally from risk at both the physiological level (34; 6) and the behavioral level" (page 4). Under this assumption, it is questionable whether forming groups based on risk preferences is a useful strategy for studying ambiguity. What do we learn about ambiguity processing when group differences are primarily based on risk preferences? Might the present results partly be driven by risk preferences rather than ambiguity preferences? I recommend the following two points: (a) Clearly justify the reasoning behind the group approach; (b) Add an additional continuous analysis approach indicating whether the key results hold independent of the group definition based on risky decision-making.

(2) k-parameter.

The authors use the k-parameter that infers the expected high-payoff probability (e.g., page 11). On page 22, this is explained as: "the subjective value term K was assigned according to each participant's internal belief of the high-payoff rate under ambiguity, yielding a participant-specific estimate of expected value under uncertainty." I hope I have not missed anything, but I neither understood the role of this parameter nor how it was computed.

(3) How were individual beliefs and models computed?

A related but more general point is that it remained unclear how the authors computed internal beliefs and internal models in the study. The study contains many statements suggesting that the authors measured internal beliefs. For example:

a) Abstract: "We show that individuals adopt distinct decision strategies that reflect different internal beliefs about unknown outcomes."
b) Page 3: "We then inferred subjective belief parameters that captured how individuals internally interpreted the ambiguous probability mass and examined how these beliefs related to choice behavior, arousal dynamics, and neural activity."
c) Page 16: "Together, these findings show that ambiguity does not evoke a uniform behavioral or physiological response across participants with different decision-making styles; instead, individuals rely on distinct internal models and computational strategies when forming decisions under ambiguity."
d) Page 16: "Taken together, these results show that ambiguity aversion is not a uniform psychological bias, but a set of heterogeneous belief-driven strategies that shape how ambiguity is represented and acted upon."
e) Page 18: "Ambiguity processing, therefore, reflects distinct belief-driven pathways rather than a single canonical mechanism."

Based on the present data, analyses, and results, I don't think that the authors can draw these conclusions. Which analyses in the manuscript identify these internal beliefs, models, or strategies? How can we dissociate a unified strategy from a heterogeneous set of strategies based on the present results? My feeling is that the k-parameter might be related to this, but as explained above, I did not understand how it was computed and what it is supposed to reflect. The DDM analyses might also be targeted at this. However, it remains elusive how the DDM captures internal beliefs about ambiguity itself. My recommendation is that the authors more clearly explain (a) why the DDM is a useful model to study ambiguity, (b) what the different parameters exactly reflect about ambiguity processing, and (c) how the DDM captures internal beliefs and distinct belief-driven strategies in this context.

(4) Statistical tests.

4.1. Figure 2B: The authors summarize the number of participants with significant effects of ambiguity on choice behavior for each group. I recommend a statistical test at the second level that properly assesses the effects of ambiguity and group within a common statistical model. In my opinion, it is not enough to simply count the number of significant tests (from the first level) for each group.

4.2. Figure 2C: For the analysis of response times, the authors might want to consider reporting the main effects of group and ambiguity.

4.3. Figure 2D: The text on page 8 states that Figure 2D indicates that "aggressive investors showed no significant pupil modulation by ambiguity...". However, the figure and its caption indicate significant differences between ambiguous and non-ambiguous trials across all groups. Moreover, if the authors want to compare the groups, it is necessary to compare the groups to each other; a test against zero within each group would not be enough to demonstrate any group differences. In my mind, this would also be important for analyses in Figure 3C and D.

4.4. Strictly speaking, for the statistical tests, it would be necessary to take into account that participants completed multiple sessions (within-subject variance is different from between-subject variance). Currently, each session is treated independently (page 19: "Each individual completed one to three experimental sessions. For data analysis, each session was treated as an independent participant, yielding a total of 108 sessions.")

(5) Necessary quality control for pupillometry and EEG data.

The task was performed in a virtual reality environment with a head-mounted display. The task was not isoluminant, and, to the best of my knowledge, participants were not instructed to avoid eye movements. The authors applied a GLM to control for luminance effects in the pupil data. For EEG, they used ICA to remove ocular and muscular artifacts. While these methods are established, they are usually applied to more controlled paradigms optimized for EEG and pupillometry. To demonstrate high data quality despite these issues, it is necessary to present quality-control analyses. Can the authors please indicate how many blinks had to be removed from the data? Could the authors please indicate how many blinks were removed from the data? Can the authors please show trial-level data (after preprocessing) for a few subjects?

(6) Quality control for the DDM.

The manuscript lacks systematic posterior predictive checks and parameter recovery for the DDM results. It is important to validate that the model accurately captures the data. Currently, we only see the model parameters, but it remains unclear whether the model performs well on the current data set. Moreover, if the authors aimed to test different strategies using the DDM, it might be useful to perform systematic model comparison.

(7) Implications of the second experiment with collaborative task remain unclear.

To me, the link between the main study and the second experiment on leadership and team performance is not obvious. In my opinion, this topic is beyond the scope of the present paper. Linking the two studies more comprehensively based on deeper theoretical grounds would likely be better suited for an independent manuscript.

Author response:

We read the Assessment as identifying two decisive gaps: (i) claims of group differences are supported by within-group tests rather than by a test of the group X condition interaction, and (ii) the drift-diffusion modeling is not yet validated or fit in a framework that permits group comparison. We agree with both, and we do not defend the current versions of these analyses. The revision will rebuild them rather than supplement them. We also agree with the reviewers that several of our conclusions about “internal beliefs” are currently stated more strongly than the analyses support, and these will be scaled back to what the modeling can carry.

Below we first note factual errors and internal inconsistencies in the manuscript that the editors asked us to flag promptly, then summarize the planned revisions, then respond to each public review comment in turn.

(1) Corrections and clarifications for the record

Reviewer #2 identified one outright error in our text, and in re-checking the manuscript we found several further inconsistencies. We list them here so that they are on the record alongside the first version of the Reviewed Preprint. In each case the reviewers' reading of the manuscript is correct and the manuscript is at fault.

(a) Aggressive investors and pupil modulation (Reviewer #2, 4.3). Our Results text states that “aggressive investors showed no significant pupil modulation by ambiguity,”. This is incorrect and contradicts our own Fig. 2d, which reports a significant, ambiguous vs. non-ambiguous difference in all three groups, including aggressive investors (T(31) = 2.74, P = 0.0304, Bonferroni-corrected). The figure and its statistics are correct; the text is wrong. The interpretive claim built on it - that aggressive investors show “blunted” arousal - is therefore unsupported and will be removed. The revised manuscript will state the correct within-group result and will test group differences directly rather than by contrasting significant against non-significant within-group tests.

(b) Within-group versus between-group claims. Relatedly, our Results state that ambiguity-related pupil differences under the subjective model are “no longer significant relative to zero,” whereas the Discussion describes them as no longer differing across groups. These are different claims, and only the first was tested. This conflation runs through several of our physiological conclusions and is the same problem the Assessment identifies. It will be resolved by replacing these statements with explicit between-group and interaction tests.

(c) Sample description (Reviewer #1, 5). The numbers are not contradictory but are certainly underspecified, and we accept that as written they cannot be reconciled by a reader. To state them plainly: 57 unique individuals each completed one to three sessions, yielding 108 sessions; 7 sessions were incomplete and 5 showed no response variability, leaving 96 sessions with usable behavior; of these, 95 had usable pupil data and 79 usable EEG. The three strategy groups are tertiles of these 96 sessions (32 each), and the smaller Ns in Figs. 2-4 (N = 31, 29, 25) reflect modality-specific exclusions within each tertile. The revision will include a participant/session flow diagram and a table reporting how many individuals contributed one, two, or three sessions, together with per-figure Ns.

(d) DDM fitting procedure (Reviewer #1, 11). One clarification: the DDMs were fit separately for each participant, not separately per group (Methods, Section 4.6); group comparisons were then performed on the participant-level parameter estimates. We note this only for accuracy of the record. The reviewer's substantive objection is unaffected and we accept it: comparing parameters estimated in independent per-participant fits does not constitute a test of group differences within a common statistical framework, and it cannot test interactions.

(e) Inconsistent specification of the EEG regressor. The trial-wise EEG covariate is described in one paragraph of Methods 4.6 as 8-13 Hz power averaged over 0-0.5 s, and in the text following Eq. 3 as a 1315 Hz difference over 0.1-0.4 s. The former corresponds to the analysis actually performed. We will correct Eq. 3's description and report the frequency band and window once, unambiguously.

(f) Time-locking and figure/caption errors. As Reviewer #1 notes (comment 14), Methods describe stimulus-locked epochs (-0.25 to 1 s) while Figs. 2-3 are labeled relative to decision onset over -1 to 2 s. We will state for each analysis whether it is stimulus- or response-locked and harmonize axis labels accordingly. In addition: the Fig. 3 caption reads "N = 25 for ideal and aggressive; N = 29 for aggressive," where the second instance should read conservative; and Section 2.6 cites Fig. 1b and 1c for the leadership and team-performance results, which are Fig. 4e and 4f.

(g) Delta/theta claims in the Discussion. Our Discussion attributes early delta- and theta-band enhancements to ideal investors. No such effects appear in our own time-frequency results, which report frontal beta-range and parietal alpha/low-beta clusters. These Discussion statements are not supported by the data presented and will be deleted. This also bears on Reviewer #1's comment 13, since it removes the only interpretation that depended on the low-frequency edge of our filter passband.

(2) Summary of planned revisions

(2.1) A single statistical framework with explicit interaction tests. All group comparisons will be replaced by unified models that include group, condition, and their interaction, with sessions nested within participants. Choice will be modeled with a generalized linear mixed model of the form invest ~ ambiguity ⨉ group + trial + (1 + ambiguity | participant/session); decision time with the corresponding linear mixed model including random slopes for ambiguity (Reviewer #1, 15). Time-resolved pupil and time-frequency EEG effects will be evaluated using cluster-based permutation tests on the group-by-condition interaction statistic rather than by aggregating within-group tests. Where our claim is that an effect is absent, most importantly, that ambiguity-related pupil differences vanish under subjective valuation, we will support it with equivalence testing and Bayes factors rather than with a non-significant P value, since a null result is not evidence for the null.

(2.2) Continuous analyses as primary; grouping justified or abandoned. We accept that trichotomizing a continuous decision tendency requires justification that we did not provide. The revision will (i) define a continuous EV-consistency index and re-run every key analysis with it as a continuous predictor, establishing that the principal conclusions do not depend on the split. The “ideal” label implies a normative optimality we have not demonstrated and will be replaced with descriptive labels (EV-consistent, EV-exceeding, EV-shortfall).

(2.3) Full specification, validation, and reliability of the belief parameter k. We agree the current description of k is inadequate. The revision will give the generative choice model in full, and explicitly state the estimation procedure and parameter bounds.

(2.4) Rebuilt and validated drift-diffusion modeling. The boundary will be freed and estimated per participant/session, so that variance is no longer forced into the drift term. We will report whether the group effect on baseline drift survives.

(2.5) Report why repeated sessions are treated as independent samples. We will conduct additional behavioral analyses to show repeated sessions from the same individual are sufficiently independent to be analyzed as unique samples. In particular, we will examine whether within-participant similarity across sessions is greater than cross-participant similarity. These analyses will provide direct evidence for whether sessions can reasonably be treated as separate samples rather than requiring sessions to be nested within participants.

(2.6) Sharper framing and appropriately scaled claims. We will remove the framing that positions the field as treating ambiguity aversion as uniform and instead situate the work within literature that already documents heterogeneity in ambiguity attitudes. The distinction between ambiguity aversion and the internal representation of ambiguity will be made explicit, with the latter identified as our actual question. Claims about “internal belief models” will be restated in terms of what is estimated. Physiological hypotheses will be stated as specific, directional, falsifiable predictions rather than by appeal to the broad range of processes EEG has been linked to.

(2.7) Quality control for pupillometry, EEG, and the VR context. We will report blink rates and counts, proportion of interpolated samples, trials and channels rejected, and ICA components removed. We will also report validate the luminance GLM.

(2.8) The collaborative task. The Apollo Distributed Control Task preceded the Lottery Choice Task by design, so it must be reported as part of the protocol regardless of the leadership analysis. However, we agree that the leadership and team-performance analyses are not sufficiently motivated to carry the interpretive weight currently given them. They will be moved to a clearly labeled exploratory section, removed from the Abstract and framing, and presented without causal or trait-level interpretation.

(3) Timeline and next steps

The revisions above require refitting the behavioral, physiological, and computational analyses rather than adding to them, including hierarchical model fitting, parameter recovery, and new quality-control analyses. We therefore anticipate submitting the revised manuscript within approximately three months, and would welcome guidance if the editors would prefer a different schedule. We are content for the first version of the Reviewed Preprint to be published with this provisional response attached.

We are grateful to both reviewers for the time invested in this manuscript. Several of the problems they identify are ones we should have caught ourselves, and the paper will be considerably stronger for their having been raised now rather than after publication.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation