Peer review process
Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.
Read more about eLife’s peer review process.Editors
- Reviewing EditorAngela LangdonNational Institute of Mental Health, Bethesda, United States of America
- Senior EditorTirin MooreStanford University, Howard Hughes Medical Institute, Stanford, United States of America
Reviewer #1 (Public review):
Summary:
The manuscript investigates value-based decision-making under risk and ambiguity using a combination of behavioral, pupillometric, and EEG data. Participants are stratified into three "decision styles" (ideal, aggressive, conservative) based on how their choices under known risk align with expected-value optimality. The central claim is that ambiguity aversion is not a uniform bias but reflects heterogeneous internal belief models, and that physiology tracks subjective belief rather than objective task structure. While this is an interesting conceptual question, the evidence is underwhelming given that differences between groups are not tested statistically (but just described), there are clear problems with how the computational models are implemented, and there are serious issues with sampling of participants.
Strengths:
The multimodal design (behavior, pupillometry, EEG) and the attempt to link a latent belief parameter to physiological signatures address a question of clear interest.
Weaknesses:
Framing and motivation
(1) The framing conflates two questions that appear distinct. The motivation centers on "ambiguity aversion," but the study's actual aim - how individuals internally represent ambiguous outcomes - seems like a different question. The relationship between these two framings needs to be made explicit, because as written the motivating phenomenon and the studied phenomenon are not obviously the same thing.
(2) Several of the contrasts the paper sets up against prior literature read as strawmen. The claim that ambiguity aversion is treated as "a single bias or fixed trait that applies uniformly" is presented as the view being overturned, but it is not clear this is a position the field actually holds - it reads as a strawman. Relatedly, the central objective-versus-subjective valuation distinction that the results are built around also reads as a strawman dichotomy rather than a genuine competing account.
(3) The motivation for the physiological measures is overly broad. The statement linking EEG to control, attention, valuation, uncertainty, conflict, effort, and engagement is so general as to be uninformative - EEG signals have been linked to essentially everything, so this does not constrain the hypotheses or predictions. A more specific, falsifiable rationale is needed.
Design, sample, and grouping.
(4) The inclusion of the collaborative spacecraft/Apollo task is unclear. It is not explained why this task is included, and its role relative to the core ambiguity question needs justification (this also bears on the leadership analyses; see below).
(5) The participant numbers do not add up and must be reconciled. The text reports 57 participants, yet the analyses describe three groups of roughly 32 + 32 + 31. The relationship between participants, sessions, and group n's needs to be stated clearly and consistently, because at present the sample description is internally contradictory.
(6) The rationale for categorizing participants into three discrete groups is not established, and the approach is statistically questionable. Decision tendency appears to be a continuous variable; dichotomizing/trichotomizing a continuous measure is generally discouraged and can manufacture or distort group differences. The authors should justify why discrete groups are needed at all, and ideally show whether there are genuine group differences (e.g., evidence of discontinuity/clustering) rather than an arbitrary split of a continuum.
Statistics
(7) Key claims about how ambiguity affects groups differently are made without the appropriate test. To support a claim that the effect of ambiguity differs across groups, the interaction (group × ambiguity) must be shown - group-wise effects reported separately are not sufficient. This is really a key limitation of the current work.
(8) The methods mentioned that some participants performed multiple sessions, but their data were treated as if coming from separate participants. This is incorrect for several reasons, particularly given the focus on individual differences.
Belief parameter and terminology
(9) The term "ideal" is not justified. It is unclear why this group is labelled "ideal" - are they Bayes-optimal, or optimal in some defined sense? If the label implies normativity, that needs to be demonstrated; otherwise it should be renamed.
Drift-diffusion modelling
(10) The boundary parameter is fixed (a detail which is hidden in the methods), but this is highly problematic. By enforcing the same boundary value for all participants, the model is forced to capture any variation as drift rate effects. As such, all conclusions about drift rate are not interpretable as they might reflect boundary effects in disguise.
(11) The DDMs are fit separately per group of participants, which again precludes testing interactions. As with the behavioral analyses, fitting separate models means group differences cannot be properly compared within a single statistical framework, and interactions cannot be assessed. The paper does mention some comparison between groups, but comparing DDM parameter estimates across separately fit models is not valid.
(12) Overall, the DDM is very complex, and the manuscript does not yet provide enough validation to make the model trustworthy. Given the number of trial-wise covariates entering the drift rate and the per-participant fitting, stronger evidence that the model is identifiable and that its parameters are recoverable/reliable is needed before the conclusions drawn from it can be accepted.
Methods - EEG and analysis details
(13) The high-pass filter setting appears very aggressive. The authors should confirm whether this risks removing genuine low-frequency signal of interest, particularly given that delta-band effects are later interpreted.
(14) There is an apparent inconsistency in the epoching/time-locking. The time-frequency analysis appears to be computed on choice-locked data, yet elsewhere the epochs are described as stimulus-locked. This needs to be clarified and made consistent, as it affects interpretation of the pre- versus post-decision EEG clusters.
(15) The mixed-effects modelling appears to omit random slopes. The authors should justify the random-effects structure (e.g., why only random intercepts), as this affects the validity of the inference.
Reviewer #2 (Public review):
Summary:
The manuscript by Qin and colleagues entitled "Pupil and Neural Dynamics Reveal Belief-Dependent Decision Making Under Ambiguity" examines decision-making under risk and ambiguity using pupillometry and EEG. The study employs a lottery choice task with three levels of ambiguity (zero, low, high). Participants were classified into three groups based on their choice behavior in a condition with risk and no ambiguity: ideal (choosing in line with objective expected values), aggressive (preference for investments), and conservative (preference against investments). The authors then compared behavior, pupil, and EEG results across these groups. The study concludes that individual beliefs about ambiguity are reflected in different behavioral strategies and neural correlates.
Strengths:
The combination of behavior, computational modeling, pupillometry, and EEG.
Weaknesses:
(1) It is unclear whether group definition is theoretically justified.
One general concern is that the strategy to form three distinct groups is not clearly motivated. The authors created the three groups, "aggressive", "ideal", and "conservative", based on the zero-ambiguity trials. However, as the authors state: "Ambiguity differs fundamentally from risk at both the physiological level (34; 6) and the behavioral level" (page 4). Under this assumption, it is questionable whether forming groups based on risk preferences is a useful strategy for studying ambiguity. What do we learn about ambiguity processing when group differences are primarily based on risk preferences? Might the present results partly be driven by risk preferences rather than ambiguity preferences? I recommend the following two points: (a) Clearly justify the reasoning behind the group approach; (b) Add an additional continuous analysis approach indicating whether the key results hold independent of the group definition based on risky decision-making.
(2) k-parameter.
The authors use the k-parameter that infers the expected high-payoff probability (e.g., page 11). On page 22, this is explained as: "the subjective value term K was assigned according to each participant's internal belief of the high-payoff rate under ambiguity, yielding a participant-specific estimate of expected value under uncertainty." I hope I have not missed anything, but I neither understood the role of this parameter nor how it was computed.
(3) How were individual beliefs and models computed?
A related but more general point is that it remained unclear how the authors computed internal beliefs and internal models in the study. The study contains many statements suggesting that the authors measured internal beliefs. For example:
a) Abstract: "We show that individuals adopt distinct decision strategies that reflect different internal beliefs about unknown outcomes."
b) Page 3: "We then inferred subjective belief parameters that captured how individuals internally interpreted the ambiguous probability mass and examined how these beliefs related to choice behavior, arousal dynamics, and neural activity."
c) Page 16: "Together, these findings show that ambiguity does not evoke a uniform behavioral or physiological response across participants with different decision-making styles; instead, individuals rely on distinct internal models and computational strategies when forming decisions under ambiguity."
d) Page 16: "Taken together, these results show that ambiguity aversion is not a uniform psychological bias, but a set of heterogeneous belief-driven strategies that shape how ambiguity is represented and acted upon."
e) Page 18: "Ambiguity processing, therefore, reflects distinct belief-driven pathways rather than a single canonical mechanism."
Based on the present data, analyses, and results, I don't think that the authors can draw these conclusions. Which analyses in the manuscript identify these internal beliefs, models, or strategies? How can we dissociate a unified strategy from a heterogeneous set of strategies based on the present results? My feeling is that the k-parameter might be related to this, but as explained above, I did not understand how it was computed and what it is supposed to reflect. The DDM analyses might also be targeted at this. However, it remains elusive how the DDM captures internal beliefs about ambiguity itself. My recommendation is that the authors more clearly explain (a) why the DDM is a useful model to study ambiguity, (b) what the different parameters exactly reflect about ambiguity processing, and (c) how the DDM captures internal beliefs and distinct belief-driven strategies in this context.
(4) Statistical tests.
4.1. Figure 2B: The authors summarize the number of participants with significant effects of ambiguity on choice behavior for each group. I recommend a statistical test at the second level that properly assesses the effects of ambiguity and group within a common statistical model. In my opinion, it is not enough to simply count the number of significant tests (from the first level) for each group.
4.2. Figure 2C: For the analysis of response times, the authors might want to consider reporting the main effects of group and ambiguity.
4.3. Figure 2D: The text on page 8 states that Figure 2D indicates that "aggressive investors showed no significant pupil modulation by ambiguity...". However, the figure and its caption indicate significant differences between ambiguous and non-ambiguous trials across all groups. Moreover, if the authors want to compare the groups, it is necessary to compare the groups to each other; a test against zero within each group would not be enough to demonstrate any group differences. In my mind, this would also be important for analyses in Figure 3C and D.
4.4. Strictly speaking, for the statistical tests, it would be necessary to take into account that participants completed multiple sessions (within-subject variance is different from between-subject variance). Currently, each session is treated independently (page 19: "Each individual completed one to three experimental sessions. For data analysis, each session was treated as an independent participant, yielding a total of 108 sessions.")
(5) Necessary quality control for pupillometry and EEG data.
The task was performed in a virtual reality environment with a head-mounted display. The task was not isoluminant, and, to the best of my knowledge, participants were not instructed to avoid eye movements. The authors applied a GLM to control for luminance effects in the pupil data. For EEG, they used ICA to remove ocular and muscular artifacts. While these methods are established, they are usually applied to more controlled paradigms optimized for EEG and pupillometry. To demonstrate high data quality despite these issues, it is necessary to present quality-control analyses. Can the authors please indicate how many blinks had to be removed from the data? Could the authors please indicate how many blinks were removed from the data? Can the authors please show trial-level data (after preprocessing) for a few subjects?
(6) Quality control for the DDM.
The manuscript lacks systematic posterior predictive checks and parameter recovery for the DDM results. It is important to validate that the model accurately captures the data. Currently, we only see the model parameters, but it remains unclear whether the model performs well on the current data set. Moreover, if the authors aimed to test different strategies using the DDM, it might be useful to perform systematic model comparison.
(7) Implications of the second experiment with collaborative task remain unclear.
To me, the link between the main study and the second experiment on leadership and team performance is not obvious. In my opinion, this topic is beyond the scope of the present paper. Linking the two studies more comprehensively based on deeper theoretical grounds would likely be better suited for an independent manuscript.