Figures and data

Overall Design.
(A) Participants, matched groups and data extracted from the ABCD dataset. (B) Domain adaptation procedure.

The gap in predictive performance (MAE) of the Baseline Model (trained only on WA) when tested on WA versus AA participants in the test set.
Confidence intervals were estimated using non-parametric bootstrap resampling (5,000 iterations) at the subject level, with resampling performed independently within AA and WA groups. For each bootstrap sample, MAEs and the gap metric were recomputed, and 95% confidence intervals were derived using the percentile method.

Adaptation Benefits.
(a) The cumulative AUIC gains of each adaptation method across all phenotypes, sorted by their performance gap scores, starting from phenotypes with the largest gap at the top left to phenotypes with the smallest gap at the bottom right. A positive AUIC means the adapted model produced a cumulative reduction in MAE as AA labels increased. (b) Repeated-measures comparison of adaptation-method AUIC gain across neuroimaging phenotypes. Each line represents an individual phenotype tracked across adaptation methods, illustrating within-phenotype changes in performance gain. The black line shows the average gain across phenotypes for each method. Colored points indicate phenotype-specific AUIC values for each method. Significance levels from Pairwise Wilcoxon signed-rank tests with Holm correction and paired Cohen’s d are shown.

(a) Changes in AA test prediction error as increasing numbers of labelled AA participants were incorporated into model training, either by direct inclusion (non-adapted) or via adaptation. Points show the mean MAE across repeated random AA subsamples at each target sample size, and shaded bands indicate approximate 95% confidence intervals across repetitions (±1.96 SD/n). Results are shown for the ten phenotypes with the largest and smallest gap scores. (b) The scatter plot of the correlation between average cumulative AUIC gain and performance gap across phenotypes.

Adaptation-specific feature re-weighting across top ten phenotypes with the (a) smallest and (b) largest performance gap.
This was estimated via a difference-of-differences framework comparing coefficient changes in adapted (Balanced Weighting) and non-adapted PLS models after adding 100 AA samples relative to the baseline (no AA added).