Mapping abstraction and metacognition onto distinct transdiagnostic symptom profiles

  1. The Department of Decoded Neurofeedback, Computational Neuroscience Laboratories, Advanced Telecommunications Research Institute International (ATR), Kyoto, Japan
  2. Developmental Computational Psychiatry, University of Tübingen, Tübingen, Germany
  3. German Centre for Mental Health (DZPG), Partner Site Tübingen, Tübingen, Germany
  4. The Department of Clinical Psychology, Graduate School of Human Sciences, Osaka University, OsakaOsaka, Japan
  5. Department of Psychology, School of Human Sciences, Senshu University, Kawasaki, Japan
  6. Graduate School of Science and Technology, Division of Information Science, Nara Institute of Science and Technology, Nara, Japan
  7. Division of Behavioural Medicine and Neurosciences, Department of Psychiatry & Behavioural Sciences, School of Medicine, Duke University, Durham, United States
  8. Department of Brain Science and Engineering, Sungkyunkwan University, Suwon, Republic of Korea
  9. Centre for Neuroscience Imaging Research, Institute for Basic Science, Suwon, Republic of Korea

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, public reviews, and a provisional response from the authors.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Nathan Faivre
    Centre National de la Recherche Scientifique, Grenoble, France
  • Senior Editor
    Jonathan Roiser
    University College London, London, United Kingdom

Reviewer #1 (Public review):

Summary:

This work investigated the associations between abstraction and metacognition in the context of reward-guided learning and transdiagnostic symptom dimensions in a sample of N = 249 participants. Participants completed a reward-guided learning task and confidence judgements. Transdiagnostic dimensions used to examine associations with abstraction and metacognition were based on a large existing dataset. Findings showed that in the examined sample, a Compulsive Hypersensitivity dimension was negatively associated with abstraction and metacognitive sensitivity, while a Social Withdrawal dimension was positively associated with metacognitive sensitivity. These data add to the existing literature on associations between metacognition and transdiagnostic symptom dimensions and extend previous work on the association between abstraction and these dimensions.

Strengths:

(1) The study addresses an interesting research question and uses a transdiagnostic approach. While part of the research question is a replication (metacognition), the additional inclusion of an abstraction parameter is highly valuable.

(2) Methodologically, the study is strong. Specifically, the implementation of an experimental task to estimate parameters, the hierarchical Bayesian modelling, the parameter recovery analyses, the bootstrap regression (including corrections for multiple comparisons) and the control of relevant covariates and response tendencies are quite impressive.

(3) The Open Science approach is laudable in general. The study was preregistered and provides open data and open code. Deviations from the preregistration are transparently reported (e.g., bootstrap regression, exclusion criteria). In this vein, the high number of robustness analyses provided in the supplements is very much appreciated.

(4) More generally, the extensiveness of the supplements is particularly valuable.

(5) Finally, it is very useful that the discussion takes into account potential alternative explanations of the findings.

Weaknesses:

(1) High number of exclusions:

As the authors mention (also in their limitations section), more than half of the participants had to be excluded (only 249 out of 512 participants remained), which is substantial. Specifically, 203 participants were excluded because they failed attention checks, 58 failed comprehension questions on the confidence scale, 25 had a reading time of the instruction page below 5 seconds, and 68 showed a performance that was too poor (the sum is probably higher than 263 because these overlap). This high number of excluded cases resulted in a substantially smaller analysed sample than planned (corresponding to approximately 77% power instead of 90%) and potentially limited both statistical power and generalisability. It might even be the case that highly impulsive individuals were excluded systematically. That is, the exclusion criteria may themselves be associated with psychiatric symptoms, so the analysed sample may no longer (fully) represent the target population. It would be interesting to see whether the results also hold when including the excluded participants (because excluded and included participants differ in attentional and impulsivity-related symptoms). A sensitivity analysis would be beneficial.

(2) Deviations from preregistration:

While the preregistration of the study is positive in general, some aspects raise doubts here. First, there have been quite a few deviations from the preregistration. For instance, the primary outcome variable for abstraction has been changed (initially: proportion of blocks in which the Abstract RL model had a better fit than the Feature RL model; instead: mean posterior responsibility of the Abstract RL model across blocks), and this appears to have influenced the findings (e.g., Supplementary Figures: 8 vs. 9). Generally, these deviations introduce additional researcher degrees of freedom and therefore warrant a more detailed justification (or replication in future studies). Second, it appears that the preregistration was uploaded only to an OSF folder and not formally registered with a timestamp. However, while an update to the document has been made according to the metadata, the content still seems to be the same as the original one (created on Jun 11, 2025).

(3) Combined measurement of choices and metacognition:

Choice behaviour and metacognition were not measured independently (i.e., they were measured using a single slider). This challenges the interpretation of results concerning metacognitive sensitivity, as the authors note themselves in the limitations section (even if the effect of starting position was small). It should be clarified whether responses exactly at the centre of the slider were possible or not (and what range the slider had).

(4) Generalisability of findings:

It is questionable whether transdiagnostic factors derived from a Japanese population can be transferred to the sample in this work, especially because it seems to be composed of a heterogeneous international sample comprising participants from multiple countries (South Africa, United States, United Kingdom, Poland, others). While it is appreciated that the authors validate the transdiagnostic factors through a new exploratory factor analysis and by correlating item loadings across samples, the correlations are only moderate and point at least to some degree of variation. In addition, for factor 2 (social withdrawal?), the correlation seems to be mainly existent due to two clusters of items, challenging the validation approach in itself. Also, apart from the correlations themselves, the absolute values of the item loadings appear to vary largely. Considering this, the conclusion that "transdiagnostic symptom dimensions appear broadly consistent across samples (even across countries)" (p. 9) definitely goes too far.

(5) Parameter recovery analysis: The parameter recovery analysis is appreciated; however, the recovery of the learning rate in particular seems to be rather low (r=.40). This raises concerns regarding the validity of the parameter estimates.

(6) Effect sizes: The reported regression coefficients appear relatively small in magnitude (i.e., betas of the associations between metacognition/abstraction and transdiagnostic dimensions appear to be in the range around .02-.04). If these are standardised estimates, the corresponding effect sizes are relatively small. If not, I recommend reporting standardised effect sizes in addition.

Overall, the study provides relevant replications and new insights into the association between reward-guided learning behaviour (abstraction, metacognition) and transdiagnostic symptom dimensions. It should be noted that effect sizes appear to be rather low, however. The analytic approach is generally strong. At the same time, the high number of exclusions, the limitations regarding the validation of the transferred transdiagnostic factor structure, and the limited recovery of the learning rate parameter clearly raise important questions regarding the robustness and generalizability of findings. Given this, any interpretations regarding potential therapeutic implications (e.g., p. 10) should be made with caution.

Reviewer #2 (Public review):

Summary:

In this study, Oka and colleagues recruited an online sample to complete a previously validated abstraction task (Cortese et al., 2021) alongside confidence ratings and a large psychiatric questionnaire battery, which included a variety of methods to screen out inattentive or otherwise biased responders. Questionnaire item scores were combined with factor weights from a large dataset to estimate transdiagnostic factor scores. A computational model was then fit in a hierarchical manner to the abstraction task data, with individuals' fit to an "Abstract RL" model used as a metric of individual-level abstraction ability, and metacognitive bias and sensitivity were estimated from the confidence ratings. Associations between these task-derived measures and both dimensional and symptom-level measures of psychopathology were then estimated using multiple regression. The key findings were that, while metacognitive sensitivity and abstraction ability were associated with symptom-level scores, the associations with transdiagnostic dimensions - higher compulsivity associated with lower abstraction ability and metacognitive sensitivity; higher social withdrawal was associated with higher metacognitive sensitivity - were interpreted by the authors as more coherent.

Strengths:

(1) Robust screening for inattentive responders through catch questions (Zorowitz et al., 2023), as well as incorporating recent recommendations regarding response bias (Sarna et al., 2026).

(2) Assessed the cross-cultural generalisability of the imported factor weights by comparing item loadings from a large external sample against a de novo exploratory factor analysis in their sample.

(3) Directly compares a theory-driven model-defined abstraction metric to metacognition in relation to dimensional and symptom-level measures of psychopathology.

(4) Pre-registered analyses, with deviations from pre-registration clearly stated.

Weaknesses:

(1) The abstraction metric (mean posterior responsibility of the Abstract RL model) differed from the pre-registered metric and has not been validated here for reliability (e.g., split-half across the blocks or similar).

(2) Model recovery is not shown, so it's not clear whether the Abstract RL and Feature RL models are fully dissociable in this task design.

(3) Metacognitive measures are behaviourally defined (AUROC2 for sensitivity and mean confidence for bias), but models do not correct for task accuracy, which may be related to both.

(4) Dimensional and symptom regressions differ: the dimensions are entered in one model, but the symptom measures are entered into separate regressions and the marginal effects corrected for multiple comparisons. If I've understood this correctly, this means that the dimensional coefficients are partial associations adjusting for the other two factors, whereas the symptom-level coefficients are marginal and FDR-corrected, making it difficult to directly compare them.

Additional questions and context:

(1) Could split-half reliability (e.g. odd vs even blocks) be reported for the abstraction measure? Relatedly, a model recovery/confusion analysis for the two abstraction models, and/or posterior predictive checks showing that the two models generate behaviour resembling that of participants would help establish that the responsibility metric is able to dissociate the different abstraction strategies.

(2) Supplementary Table 2 shows the results for the pre-registered discrete proportion metric - here, there is limited evidence (p=0.220) of an association between abstraction and compulsivity, so saying they are "almost consistent" is perhaps a little overstated. Though the argument for using the alternative continuous metric is justified in the text, it's not quite clear whether the difference is due to the inference method or the abstraction metric itself - the bootstrapped analysis of the pre-registered metric is not reported, nor is the analysis without bootstrapping of the continuous metric (I think this may have been what Supplementary Table 1 was meant to report, but currently it's identical to Supplementary Table 5). In addition, it might be helpful if the correlation between the two metrics were presented graphically.

(3) In the Methods and Supplement, the authors mention that they had pre-registered running a sensitivity analysis including excluded participants. This might be interesting given the high exclusion rate, and given that most exclusions were not based on task behaviour (chance-level choosing). If there is concern about shifting group-level parameter distributions, then this could be explicitly included in the model by including an offset on group-level parameters (i.e., interaction term) on excluded participants, which would allow them to systematically differ in model parameters. Alternatively, one could at least estimate the abstraction and metacognition metrics in the excluded sample (perhaps restricted to those excluded on questionnaire-based criteria rather than task performance) to see whether they do indeed differ.

(4) How do factor scores relate to task accuracy - do those with higher compulsivity perform worse, and is this plausibly related to less abstraction?

(5) How do the factors extracted here compare to those in other studies, such as those from Gillan et al. (2016, eLife)? In particular, it'd be interesting to know what questionnaires/symptoms in the "Compulsive hypersensitivity" factor in the present study overlap with the Compulsive behaviour/intrusive thoughts factor from that earlier work, as the latter has been strongly associated with metacognitive measures - higher metacognitive efficiency for anxious/depression, lower metacognitive efficiency for compulsive behaviour - in previous work (Rouault et al., 2018). That three-factor structure also included a factor they labelled "Social withdrawal" - is it similar to the one presented here, or is the one here (including distress) more like their anxious depressive factor?

(6) In the Discussion, the authors state "Our findings are also consistent with previous converging evidence linking compulsive tendencies to less efficient computation and a preference for familiar over goal-directed action". That the Feature RL model might fairly be called less efficient is reasonable, but I'm not sure how the reduction of features in the Abstract RL model relates to goal-directed action (they're both model-free RL algorithms).

(7) The Discussion also mentions "models that integrate abstract and metacognitive representations" - was there a reason these could not be applied in the present study?

Reviewer #3 (Public review):

Summary:

This is an interesting study investigating the relationship between transdiagnostic symptom profiles and computational parameters of abstraction and metacognition. The authors found that a transdiagnostic dimension they term compulsive hypersensitivity was negatively related to abstraction ability and metacognitive sensitivity, and a dimension termed social withdrawal was positively related to metacognitive sensitivity. Overall, the question of whether and how higher-level cognitive processes relate to transdiagnostic psychiatric factors is interesting and addresses a relevant gap in the literature.

Strengths:

This is an overall well-designed study, and the authors have clearly given considerable thought to data quality and validation during the study design. This is evident, for instance, in the use of multiple approaches to assess inattentive responding and acquiescence tendencies. In addition, the use of advanced modelling approaches for the task data represents a strength and allows the authors to capture individual differences in task behaviour in a more nuanced and mechanistic manner in comparison to what would have been possible with model-agnostic analyses.

Weaknesses:

Nevertheless, I would like to raise several concerns, detailed below.

(1) My most pressing concern regards the exclusion rate of roughly half the sample. Of the 512 participants who were tested, only 249 (48.6%) were included in the final analyses, meaning that more than half of the recruited participants were excluded. Although the authors state that these exclusions were based on preregistered criteria, this high exclusion rate still warrants further investigation. Pre-registering exclusion criteria does not eliminate the potential for selection bias or establish that the resulting sample is representative of the recruited sample. In fact, the authors even report that the excluded participants differed significantly on several psychiatric scores. I believe that a detailed account of the exclusion, including how included and excluded participants differed on relevant demographic or study variables, would be beneficial.

Additionally, the authors state that a sensitivity analysis including participants excluded from the primary analysis was preregistered but was not conducted because the excluded and included participants differed significantly. This does not seem a sufficient justification for omitting a preregistered sensitivity analysis; in fact, systematic differences between included and excluded participants warrant this kind of sensitivity analysis. I agree with the authors that this may lead to changes in task parameter estimates. However, I believe this change would be an informative result of the sensitivity analyses rather than a methodological problem to avoid.

Finally, I would like the authors to clarify some inconsistencies in the preregistered exclusion criteria. In the pre-registration, they first state that all participants who fail any infrequency question will be excluded. Later in the pre-registration, they state that participants with two or more failed questions will be excluded and that a sensitivity analysis will be done, in which participants with only one failed question are retained. Neither of these criteria matches what is reported in the manuscript ("Attention check + infrequency item mistake more than 2").

(2) The current sample size of n = 249 participants is not sufficient for a factor analysis with 176 questionnaire items, and I believe that the solution of using factor weights from a previous larger sample is generally sensible. I also commend the authors for wanting to assess the validity of this approach. However, both the methods and results reported for the comparison between the factor analysis based on the current, smaller sample and the large, previous sample are insufficient and warrant substantially more detail. Firstly, it is unclear to me whether the three-factor solution in the factor analysis with the current, smaller sample was empirically grounded or whether three factors were extracted to match the three-factor solution from the larger sample. For both factor analyses, I would recommend reporting the factor extraction methods, the criterion used to determine the number of factors, and whether alternative factor solutions were considered. Secondly, it is unclear what exactly was correlated across the two factor analyses (e.g., factor scores, item loadings, factor weights?). Finally, I do not believe that the result of significant correlation is sufficient to conclude that the dimensions replicate between samples, particularly given they indicate only moderate correspondence (r = 0.46, for instance, corresponds to only 21% shared variance), thereby providing very limited evidence for replication of the factor structure.

These concerns are particularly pressing given that the authors state the consistency of transdiagnostic symptom dimensions across samples and countries as a primary result in the discussion.

(3) It is also unclear to me whether any exclusion was applied based on the first part of the Deary-Liewald reaction time task. The supplement implies that it was ("This task result was used to [...] exclude participants who exhibit problematic behaviours"), but I could not find a corresponding report in the manuscript.

(4) The manuscript refers to an attentional check questionnaire used to identify and exclude inattentive participants, but does not provide a reference for this measure or describe the questions included. Given the substantial overall exclusion rate, it is particularly important that all exclusion procedures and criteria are described in sufficient detail to allow readers to assess their appropriateness and reproducibility. The authors should therefore provide the relevant reference and/or report the specific questions and criteria used to determine inattention.

Relatedly, when describing the control-neutral items in the Supplementary Materials, the authors state that further details are provided in the Supplementary Materials. However, I cannot find another, separate Supplementary Material that may contain this information

Author response:

We thank the editors for the eLife Assessment and the reviewers for their thorough and constructive evaluation of our manuscript.

We are glad that the evidence for reduced metacognitive sensitivity in relation to the compulsive hypersensitivity dimension was considered solid. In hindsight, we agree that the evidence for the abstraction findings is currently incomplete, and we plan to address this with additional analyses as outlined below (to clarify the robustness of the abstraction metric and its association with symptom dimensions).

Below, we outline how we plan to address the main points raised by the reviewers, grouped thematically, given the overlap across the three reviews.

A full, detailed point-by-point response accompanied by the corresponding analyses will follow in our formal revision response.

(1) Exclusion rate and lack of relevant sensitivity analysis

In the revision, we will report a detailed comparison of included versus excluded participants on demographic and symptom variables, and we will conduct the originally preregistered sensitivity analysis to estimate abstraction and metacognition metrics, including the excluded sample, and compare these metrics with psychopathological variables, rather than omitting this analysis.

(2) Abstraction metric and its deviation from the preregistered definition

We recognise that our primary measure of abstraction differs from the preregistered metric (the proportion of blocks better fit by the Abstract RL model), and that this change appears to affect the pattern of results. In the revision, we will report split-half reliability for the abstraction metric, present the bootstrapped analysis of the preregistered metric, and provide a clearer justification for the switch, making the source of the discrepancy (metric versus inference method) transparent.

(3) Generalisability of the transdiagnostic factor structure

We acknowledge that our claim of consistency between samples of the transdiagnostic dimensions requires more support and more careful framing. In the revision, we will provide full methodological detail on both factor analyses (extraction method, criteria for the number of factors, etc.), and we will moderate our interpretation, especially in the discussion, to reflect the actual strength of correspondence rather than describing the structure as broadly consistent.

(4) Parameter and model recovery

We will extend our recovery analyses to include a model recovery/confusion analysis between the Abstract RL and Feature RL models, and we will investigate the source of the relatively low learning-rate recovery (e.g., by testing recovery stratified by block length and without injected noise), reporting these results in the revision.

(5) Documentation of the attention-check procedure

We will provide all details on the attention-check items and criteria used to identify inattentive participants, including a reference where applicable, and we will standardise terminology (e.g., infrequency vs. inattention items) throughout the manuscript and supplementary information.

(6) Other methodological and presentational clarifications

We will review the manuscript, methods, and supplementary materials to resolve the remaining methodological ambiguities and presentational issues raised by the reviewers. This includes, among other points, clarifying whether the reported regression coefficients are standardised and correcting labelling inconsistencies, missing references/DOIs, and other minor textual issues throughout.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation