Comprehensive characterization of human color discrimination thresholds
Figures
Task, stimuli, and the Wishart process psychophysical model (WPPM).
(A) 3AFC oddity task. On each trial, participants viewed a triplet of stimuli—two identical references and one different comparison—and identified the odd one out. (B) Stimuli are constrained to lie in the isoluminant plane that passes through the monitor’s gray point. Data are represented and fit in a transformation of this plane, which we refer to as , a square bounded between −1 and 1. The grid of dots illustrates the transformation between the plane in RGB space and model space. (C) Example of a smoothly varying covariance matrix field produced by the WPPM. The field is generated by sampling from a smooth finite-basis Wishart process prior ( and ; see more details in Prior over the weight matrix). Although the field is illustrated on a 7 × 7 grid, it specifies a covariance matrix for every stimulus in the plane, as shown in the heatmaps. (D) Example of a less smoothly varying covariance matrix field. This field is obtained by sampling from a less smooth Wishart process prior ( and ). (E) Observer model. For each stimulus triplet , internal representations are drawn from multivariate Gaussian distributions, each centered on its corresponding stimulus, with noise characterized by its corresponding covariance matrix. The model determines whether the observer correctly identifies the odd stimulus by comparing the squared Mahalanobis distances between all three pairs. (F) Derivation of the elliptical threshold contour. One-dimensional psychometric functions are approximated using Monte Carlo simulations (10,000 samples per stimulus pair shown for illustration; 2000 used during model fitting). For each selected chromatic direction, we derive the threshold point corresponding to 66.7% correct. An ellipse is then fit to the resulting threshold points to describe the discrimination threshold contour.
Threshold results and validation.
(A) Adaptively sampled trials. AEPsych-driven stimulus pairs are sampled to estimate thresholds across the psychometric field. Of the 6000 trials, the first 900 are Sobol'-sampled; the remaining 5100 (shown here) are adaptively selected using the expected absolute volume change (EAVC) acquisition function, based on a nonparametric Gaussian process (GP) model updated every 20 trials. (B) Discrimination threshold contours (66.7% correct) read out from the Wishart process psychophysical model (WPPM) on a grid of reference stimuli for a representative participant, based on fits to the 6000 AEPsych trials. (C) Group summary of WPPM readouts (), evaluated on the same grid of reference stimuli. (D) Validation trials for the same participant. The validation conditions (reference stimuli and chromatic directions along which the comparison stimulus varies) are randomly generated for each participant (see Appendix 4 for validation conditions used for the remaining participants). (E) Comparison of thresholds. Ellipses represent discrimination threshold contours read out from the WPPM fit (same fit as in panel B), evaluated at the 25 reference stimuli used in the validation trials. Gray lines: the validation directions; black bars: the 95% bootstrapped confidence intervals for the corresponding validation thresholds. (F) Comparison of psychometric functions. Only two validation conditions are shown for illustration (see Appendix 4.1 for all 25 conditions for each participant). (G) Linear regression of thresholds predicted by the WPPM against validation thresholds for the same participant. Horizontal and vertical error bars represent 95% confidence intervals for the validation thresholds and WPPM predictions, respectively. (H) Summary of regression slopes and correlation coefficients for all participants. Error bars: 95% confidence intervals. As a benchmark, the same analysis is performed on a dataset simulated using a ground-truth WPPM instance that approximates CIELAB ΔE94 (Appendix 5).
Comparison of color discrimination thresholds with previous measurements.
Across all panels, black contours represent thresholds from prior studies, whereas colored contours represent the 66.7% discrimination thresholds estimated in our study. Shaded regions indicate 95% confidence intervals from 120 bootstrapped datasets. (A) MacAdam, 1942. Top panel: MacAdam’s original threshold ellipses, magnified 10× for visualization. Bottom panel: Threshold contours measured from one participant in our study and transformed from model space into the CIE 1931 chromaticity diagram. Reference stimuli are sampled from a 5 × 5 grid spanning from –0.7 to 0.7 along each dimension of model space. To reduce visual clutter, MacAdam ellipses falling within the gamut of our isoluminant plane (parallelogram) are shown only by arrows indicating their major axes. For visual comparability, our ellipses are magnified 2× to approximately match the scale of MacAdam’s data. The triangle indicates our monitor gamut. (B) Danilova and Mollon, 2025. Left panel: Original threshold contours (79.4% correct) from their study, magnified by 4×. Right panel: Threshold contours from one participant in our study, transformed from model space into the scaled MacLeod–Boynton space used in their study. Reference points are sampled on a 5 × 5 grid ranging from –0.7 to 0.7. As in panel A, to reduce visual clutter, ellipses from their study that fall within the gamut of our isoluminant plane (parallelogram) are shown as black arrows indicating only their major axes. For visual comparability, our ellipses are magnified 1.5×. (C) Krauskopf and Gegenfurtner, 1992. Left panel: 79.4% threshold contours. Right panel: 66.7% threshold contours from one participant in our study, transformed into DKL space with the axes scaled for each participant to normalize the thresholds along the L–M and S axes at the adapting chromaticity. All contours are shown at their original sizes in this scaled representation. (D) CIELAB ΔE76, ΔE94, and ΔE00. Threshold is defined as , chosen to approximately match the scale of our measured thresholds, which are shown at their original sizes. See Appendices 6–9 for additional details.
© 1992, Krauskopf and Gegenfurtner. Figure 3C left hand image is reproduced from Figure 14 from Krauskopf and Gegenfurtner, 1992 (published under a CC-BY-NC-ND). Further reproductions must adhere to the terms of this license.
The finite-basis Wishart process psychophysical model (WPPM).
(A) Model overview. In our implementation, we use a set of 5 × 5 two-dimensional Chebyshev polynomial basis functions, denoted , where . These basis functions are combined using a learnable weight matrix W to produce an overcomplete representation , where and . The resulting representation is then combined with its transpose to produce a field of symmetric positive semi-definite covariance matrices. Each matrix specifies internal noise in terms of the variance along the two model dimensions, and , and their covariance, . The example covariance matrix field shown here is generated from the best-fitting weights for participant CH (see Appendix 2—figure 3 for all participants). (B) Model readouts. Internal noise can be read out anywhere in model space, illustrated here on a 7 × 7 grid of reference stimuli (solid lines), from which threshold contours (dashed lines) can be derived.
AEPsych-driven trials (900 Sobol'-sampled and 5100 adaptively sampled), fallback Sobol'-sampled trials, and Wishart process psychophysical model (WPPM) predictions for all participants.
Each row shows data from one participant.
Percent correct as a function of the angular difference between the reference and comparison stimuli for all participants.
The number of trials within each bin (bin width = 4°) is encoded by both marker size and color, with larger markers and brighter colors indicating more trials.
Covariance matrix fields obtained from the best-fitting Wishart process psychophysical model (WPPM) for all participants.
Rows show, from top to bottom, the noise variance along the first model dimension, the noise variance along the second model dimension, and the covariance between the two dimensions.
Best-fitting weights for all participants.
The symmetric solid curves show the prior imposed on the model. The prior was implemented by specifying the variance of the weights, η, as a function of the polynomial order of the Chebyshev basis functions and two hyperparameters, ε and γ (Equation 16). The solid curves indicate for our chosen hyperparameters, corresponding to two standard deviations of the prior distribution. The horizontal positions of the dots are jittered to improve visibility.
Task timing and real-time trial scheduling.
(A) Trial sequence: a 0.5 s fixation cross was followed by a 0.2 s blank interval, and then a 1 s presentation of three blobby stimuli. Participants responded at their own pace to identify the odd one out, after which a 0.2 s blank screen and 0.5 s feedback were shown. The inter-trial interval (ITI) was 1.5 s. (B) Schematic representation of the trial timing and computational responsibilities of the two computers.
Validation for participant ME.
Same format as Figure 2D–G in the main text.
Validation for participant SG.
Same format as Figure 2D–G in the main text.
Validation for participant DK.
Same format as Figure 2D–G in the main text.
Validation for participant BH.
Same format as Figure 2D–G in the main text.
Validation for participant FM.
Same format as Figure 2D–G in the main text.
Validation for participant HG.
Same format as Figure 2D–G in the main text.
Validation for participant FW.
Same format as Figure 2D–G in the main text.
Validation for participant CH.
Same format as Figure 2D–G in the main text.
Threshold residuals.
Data are pooled across all validation conditions and all participants (). In all panels, color indicates the surface color of the reference stimulus, and the y-axis limits are set to ± the mean validation threshold. (A) Residuals as a function of the absolute angular difference between the major axis of the elliptical threshold contours read out from the Wishart process psychophysical model (WPPM) fits and the chromatic direction of the validation condition. (B) Residuals as a function of the aspect ratio (major/minor axis) of the WPPM threshold contours. (C) Residuals as a function of thresholds estimated from the validation trials.
Derivation of the ground-truth Wishart process psychophysical model (WPPM) fit based on CIELAB ΔE94.
(A, B) Comparison stimuli at the iso-distance contours in the isoluminant plane, shown in both RGB and model spaces. Note that the reference grid and fixed set of directions shown here are for illustration only; the actual sampling did not use a fixed grid or evenly spaced chromatic directions. (C) The Weibull psychometric function used to simulate binary (correct or incorrect) responses given ΔE values. (D) Sampled reference–comparison stimulus pairs. Reference colors and chromatic directions were sampled using Sobol' sequences, and comparison stimuli were jittered around the iso-distance contour. A total of 18,000 trials were simulated; only the first 200 are shown here for clarity. (E) Comparison between readouts from the WPPM fit and from CIELAB ΔE94. The WPPM fit was subsequently treated as the ground truth when simulating AEPsych and validation trials.
AEPsych-driven trials and Wishart process psychophysical model (WPPM) readouts for a simulated observer.
(A) AEPsych-driven Sobol’-sampled trials. (B) AEPsych-driven adaptively sampled trials. (C) Comparison of WPPM predictions with the ground truth, shown here as dashed contours and in Appendix 5—figure 1E as colored contours.
Validation trials and Wishart process psychophysical model (WPPM) readouts for a simulated observer.
Same format as Figure 2D–G in the main text.
Threshold residuals for a simulated dataset.
For all panels, color indicates the surface color of the reference stimulus, and the y-axis limits are set to ± the mean of the validation thresholds. (A) Residuals as a function of the absolute angular difference between the major axis of the elliptical threshold contours read out from the Wishart process psychophysical model (WPPM) fits and the chromatic direction of the validation condition. (B) Residuals as a function of the aspect ratio (major/minor axis) of the WPPM threshold contours. (C) Residuals as a function of thresholds estimated from validation trials.
Deviation of Wishart process psychophysical model (WPPM) estimates from the ground truth.
(A) Bures–Wasserstein (BW) distance between WPPM-estimated thresholds and the ground-truth ellipses. The upper limit of the color map (0.17) corresponds to the maximum BW distance between each ground-truth ellipse and a reference circle whose radius equals the largest major axis length among all ground-truth ellipses. The maximum BW distance between WPPM estimates and the ground truth (0.03) is substantially lower than this reference value. (B) Difference in major-axis length between WPPM readouts and ground-truth ellipses. The colormap limits (±0.17) reflect the ± maximum ground-truth major axis length. Again, the maximum deviation observed (0.03) is small relative to this range.
Comparison with MacAdam, 1942.
First panel: MacAdam’s original threshold ellipses, magnified 10× for visualization. Remaining panels: threshold contours corresponding to 66.7% correct (colored lines), measured from all participants and transformed from model space into the CIE 1931 chromaticity diagram. Shaded regions indicate 95% confidence intervals computed from 120 bootstrapped datasets. Reference stimuli were sampled from a 5 × 5 grid spanning [–0.7, 0.7] along each model dimension. To reduce visual clutter, MacAdam ellipses falling within the gamut of the isoluminant plane are represented only by arrows indicating their major axes. For visual comparability, our ellipses are magnified 2× to approximately match the scale of MacAdam’s data. Triangle: monitor gamut; quadrilateral: gamut of the isoluminant plane.
Comparison with Danilova and Mollon, 2025, in the scaled MacLeod–Boynton space.
Top left: threshold contours from their study (black ellipses), magnified by 4×. Remaining panels: threshold contours from all participants (colored ellipses). We sampled a grid of reference points evenly spaced from –0.7 to 0.7 in our model space, read out the corresponding threshold contours, and transformed them into the same scaled MacLeod–Boynton space. The parallelogram indicates the gamut of the isoluminant plane. To reduce visual clutter, ellipses from their study that fall within our gamut are represented by arrows indicating only their major axes. For visual comparability, our ellipses are magnified by 1.5× to roughly match the size of those in their study.
Transformation from model space to the stretched DKL space used in Krauskopf and Gegenfurtner, 1992, for participant CH.
(A) Model space. Threshold contours corresponding to 66.7% correct were read out from each participant’s Wishart process psychophysical model (WPPM) fit. Notably, our measurements covered a much larger region of the isoluminant plane than did theirs. (B) Intermediate, unstretched DKL space, obtained by affine transformation from model space. (C) Stretched DKL space, obtained by affine transformation from unstretched DKL space. Specifically, the cardinal axes of unstretched DKL space were rescaled to normalize the threshold at the achromatic reference point.
Comparison with Krauskopf and Gegenfurtner, 1992, for the remaining seven participants.
Top left: original threshold contours reported by Krauskopf and Gegenfurtner, 1992. Remaining panels: threshold contours (colored lines) for the remaining participants, transformed from model space to stretched DKL space using participant-specific scaling of the cardinal axes. Shaded regions indicate 95% confidence intervals computed from 120 bootstrapped datasets. All threshold contours are plotted at their original sizes.
© 1992, Krauskopf and Gegenfurtner. Top-left image is reproduced from Figure 14 from Krauskopf and Gegenfurtner, 1992 (published under a CC-BY-NC-ND). Further reproductions must adhere to the terms of this license.
Comparison with color differences predicted by CIELAB ΔE76 (CIE, 2004).
The CIELAB threshold was defined as , chosen to approximately match the scale of the measured thresholds in our data. Black contours represent the CIELAB predictions, whereas colored contours represent the measured thresholds transformed from model space into L*a*b* space and shown at their original scales. Shaded regions indicate 95% confidence intervals computed from 120 bootstrapped datasets.
Comparison with predictions based on the CIELAB ΔE94 color-difference metric (CIE, 1995).
Comparison with CIELAB ΔE00 color-difference metric (Luo et al., 2001; CIE, 2001).
The effects of ε and γ on the variance of model weights.
(A) Variance of the Chebyshev basis weights as a function of polynomial order (). The top panel illustrates effects of varying ε while holding γ fixed, whereas the bottom panel shows effects of varying γ while holding ε fixed. The yellow dashed curve indicates the hyperparameter values used in the main analyses (, ). (B) Two-dimensional Chebyshev basis functions arranged in order of increasing polynomial order ().
Effect of ε on the Wishart process psychophysical model (WPPM)-predicted psychometric field for a representative participant.
The gray line and shaded region indicate the mean and full range of negative log-likelihood (nLL) on the training set across five repetitions of fivefold cross-validation. The green line and shaded region indicate the mean and full range of nLL on the test set. Panels (a–g) show the model-predicted thresholds on a 7 × 7 reference grid for selected values of ε.
Effect of γ on the Wishart process psychophysical model (WPPM)-predicted psychometric field for a representative participant.
The gray line and shaded region indicate the mean and full range of negative log-likelihood (nLL) on the training set across five repetitions of fivefold cross-validation. The green line and shaded region indicate the mean and full range of the nLL on the test set. Panels (a–g) show the model-predicted thresholds on a 7 × 7 reference grid for selected values of γ.
Effect of ε on the slope and correlation coefficient of the linear regression between Wishart process psychophysical model (WPPM)-predicted and validation thresholds.
Stimuli and equipment used for calibration.
(A) The stimulus setup during calibration was identical to that used in the main experiment. The surface color of both the cubic room and the blobby stimulus (shown here as the top-position stimulus) was varied during calibration. The shaded gray circular region on the stimulus indicates the area measured by the spectroradiometer lens. (B) A SpectraScan PR-670 used for all calibration measurements.
Characterization of display output through Unity’s rendering pipeline.
(A) Gamma functions for the red, green, and blue primaries. Note that they lie above the identity line with Unity’s default internal gamma correction. (B) Spectral power distributions (SPDs) of the three primaries across a range of intensity levels. (C) Chromaticities of the three primaries at different intensity levels. (D) Normalized SPDs for each primary, showing that SPD shape is stable across intensity levels. (E) Linearity tests comparing predicted and measured chromaticity and luminance across two independent measurement runs. (F) Prediction errors, computed as the differences between measured and nominal (predicted) values, showing little systematic deviation from zero. (G) Effect of the cubic room’s background color on the SPD of the blobby stimulus, showing no detectable influence.
Comparison of display output across stimulus locations.
(A) Spectral power distributions (SPDs) for each stimulus location: Ref Cal (bottom right), Cal 2 (bottom left), and Cal 3 (top). (B) Ambient light SPDs measured during calibration. (C) Gamma functions for each primary (red, green, blue) across all three stimulus locations. (D) Differences in normalized output for each pairwise comparison of stimulus locations, plotted separately for each primary. (E) Chromaticity coordinates of each primary in the CIE diagram, shown for all three stimulus locations.
Gamma correction.
(A) Measured gamma functions and corresponding inverse functions for the red, green, and blue primaries, used to construct the gamma-correction lookup table. (B) Gamma functions remeasured after applying gamma correction in Unity, showing close alignment with the identity line for all three primaries.
Comparison of display output with gamma correction over time.
(A) Spectral power distributions (SPDs) measured at the bottom-right blobby stimulus location. Ref Cal denotes the initial measurements before the experiment, and Cal 2 denotes the follow-up measurements conducted roughly 1 month after data collection began. (B) Ambient-light SPDs measured during each calibration. (C) Gamma functions for the red, green, and blue primaries across both sessions, with gamma correction applied. (D) Chromaticity coordinates of each primary plotted on the CIE chromaticity diagram for both calibration runs.
Evidence of spatial dithering by Unity’s rendering pipeline.
(A) The stimulus setup during measurement was similar to that used in the main experiment, except that only a single blobby stimulus was presented at the center of the screen and the cubic room was omitted. The shaded gray circular region on the stimulus indicates the area measured by the colorimeter aperture. (B) Spatial dithering by Unity’s standard shader is suggested by comparing luminance measurements from the Klein K-10A (averaged across a circular region on the blobby object) with the RGB values stored in the frame buffer. The measured luminance shows small incremental changes as the RGB settings increase in steps of 1/4095. These measurements are consistent with the values obtained by averaging over pixels in saved frame-buffer images (exported from Unity in .exr format). The averaged pixel values exhibit 12-bit quantization, even though individual pixel values exhibit 8-bit quantization. (C) Top row: mean R channel values averaged vertically within a horizontal slice of the blobby object. Bottom row: differences in the R channel values between the minimum target R channel setting and each of the remaining settings. Different shades of gray represent different target R settings. For illustration, only a portion of the horizontal slice is shown, and solid lines in the bottom row are scaled by a factor of 0.1. Dashed lines: mean difference averaged across all pixels within each slice.
Tables
Corner vertices in DKL, LMS, RGB, and model spaces.
| Corner | DKLL-M | DKLS | DKLLum | L | M | S | R | G | B | Wdim1 | Wdim2 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | –0.123 | –0.812 | 0 | 0.145 | 0.147 | 0.016 | 0.000 | 0.733 | 0.000 | –1 | –1 |
| 2 | 0.175 | –0.830 | 0 | 0.164 | 0.111 | 0.014 | 1.000 | 0.407 | 0.000 | 1 | –1 |
| 3 | –0.175 | 0.830 | 0 | 0.142 | 0.154 | 0.152 | 0.000 | 0.593 | 1.000 | –1 | 1 |
| 4 | 0.123 | 0.812 | 0 | 0.160 | 0.117 | 0.150 | 1.000 | 0.267 | 1.000 | 1 | 1 |
Transformation matrices between DKL, RGB, and model spaces.
| MDKL→W | MRGB→W |
|---|---|
Linear regression results assessing the relationship between WPPM–validation threshold residuals and three predictors: (1) the absolute angular difference between the chromatic direction of the validation condition and the major axis of the threshold contours read out from the WPPM fits, (2) the aspect ratio of the threshold contours, and (3) the magnitude of the validation threshold.
| Predictor | Term | Coef | Std Err | t | p | [0.025, 0.975] CI | R2 |
|---|---|---|---|---|---|---|---|
| Absolute difference of angles | Intercept | 0.004 | 0.002 | 2.254 | 0.025 | [0.000, 0.007] | 0.000 |
| Slope | 0.000 | 0.000 | 0.273 | 0.785 | [–0.000, 0.000] | ||
| Aspect ratio | Intercept | 0.001 | 0.004 | 0.319 | 0.750 | [–0.006, 0.008] | 0.004 |
| Slope | 0.002 | 0.002 | 0.943 | 0.347 | [–0.002, 0.005] | ||
| Validation thresholds | Intercept | 0.014 | 0.002 | 7.511 | 0.000 | [0.010, 0.018] | 0.142 |
| Slope | –0.176 | 0.031 | –5.727 | 0.000 | [–0.237,–0.116] |
Catch trial performance summary across all sessions.
Lower and upper bounds indicate the participant’s lowest and highest session-level performance, respectively.
| Participant | ME | SG | DK | BH | FM | HG | FW | CH |
|---|---|---|---|---|---|---|---|---|
| Proportion correct | 0.996 | 0.996 | 0.948 | 0.992 | 0.998 | 0.988 | 0.988 | 0.998 |
| Lower bound | 0.977 | 0.974 | 0.868 | 0.975 | 0.975 | 0.954 | 0.957 | 0.980 |
| Upper bound | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 | 1.000 |
Linear regression results for the simulated dataset.
| Predictor | Term | Coef | Std Err | t | p | [0.025, 0.975] CI | R2 |
|---|---|---|---|---|---|---|---|
| Absolute difference of angles | Intercept | –0.007 | 0.005 | –1.532 | 0.139 | [–0.017, 0.003] | 0.028 |
| Slope | 0.000 | 0.000 | 0.812 | 0.425 | [–0.000, 0.000] | ||
| Aspect ratio | Intercept | –0.007 | 0.012 | –0.579 | 0.568 | [–0.033, 0.019] | 0.003 |
| Slope | 0.002 | 0.008 | 0.263 | 0.795 | [–0.014, 0.018] | ||
| Validation thresholds | Intercept | 0.028 | 0.006 | 4.429 | 0.000 | [0.015, 0.041] | 0.547 |
| Slope | –0.393 | 0.075 | –5.273 | 0.000 | [–0.547, –0.239] |
Regression slopes assessing the relationship between Wishart process psychophysical model (WPPM)–validation threshold residuals and the three predictors reported in Appendix 4—table 1.
The hyperparameter γ was fixed at 0.0003, while ε was varied.
| Predictor | ε | Coef | Std Err | t | p | [0.025, 0.975] CI | R2 |
|---|---|---|---|---|---|---|---|
| Absolute difference of angles | 0.1 | –0.001 | 0.000 | –6.926 | <0.001 | [–0.001, –0.000] | 0.195 |
| 0.2 | –0.000 | 0.000 | –3.212 | 0.002 | [–0.000, –0.000] | 0.050 | |
| 0.3 | 0.000 | 0.000 | 1.764 | 0.079 | [0.000, 0.000] | 0.015 | |
| 0.4 | 0.000 | 0.000 | 0.273 | 0.785 | [–0.000, 0.000] | 0.000 | |
| 0.5 | –0.000 | 0.000 | –0.654 | 0.514 | [–0.000, 0.000] | 0.002 | |
| 0.6 | –0.000 | 0.000 | –1.393 | 0.165 | [–0.000, 0.000] | 0.010 | |
| 0.7 | –0.000 | 0.000 | –1.592 | 0.113 | [–0.000, 0.000] | 0.013 | |
| 0.8 | –0.000 | 0.000 | –1.901 | 0.059 | [–0.000, 0.000] | 0.018 | |
| 0.9 | –0.000 | 0.000 | –3.732 | <0.001 | [–0.000, –0.000] | 0.066 | |
| 1.0 | –0.000 | 0.000 | –1.961 | 0.051 | [–0.000, 0.000] | 0.019 | |
| Aspect ratio | 0.1 | –0.003 | 0.001 | –2.542 | 0.012 | [–0.006, –0.001] | 0.032 |
| 0.2 | 0.001 | 0.005 | 0.282 | 0.778 | [–0.008, 0.011] | 0.000 | |
| 0.3 | 0.003 | 0.002 | 1.068 | 0.287 | [–0.002, 0.007] | 0.006 | |
| 0.4 | 0.002 | 0.002 | 0.943 | 0.347 | [–0.002, 0.005] | 0.004 | |
| 0.5 | 0.001 | 0.002 | 0.289 | 0.773 | [–0.003, 0.004] | 0.000 | |
| 0.6 | –0.001 | 0.002 | –0.724 | 0.470 | [–0.004, 0.002] | 0.003 | |
| 0.7 | –0.004 | 0.001 | –2.625 | 0.009 | [–0.006, –0.001] | 0.034 | |
| 0.8 | –0.005 | 0.001 | –4.596 | <0.001 | [–0.008, –0.003] | 0.096 | |
| 0.9 | –0.002 | 0.001 | –3.266 | 0.001 | [–0.003, –0.001] | 0.051 | |
| 1.0 | –0.002 | 0.001 | –3.141 | 0.002 | [–0.003, –0.001] | 0.047 | |
| Validation thresholds | 0.1 | –0.964 | 0.061 | –15.767 | <0.001 | [–1.085, –0.844] | 0.557 |
| 0.2 | –0.417 | 0.049 | –8.580 | <0.001 | [–0.513, –0.321] | 0.271 | |
| 0.3 | –0.266 | 0.032 | –8.245 | <0.001 | [–0.329, –0.202] | 0.256 | |
| 0.4 | –0.176 | 0.031 | –5.727 | <0.001 | [–0.237, –0.116] | 0.142 | |
| 0.5 | –0.144 | 0.034 | –4.207 | <0.001 | [–0.211, –0.076] | 0.082 | |
| 0.6 | –0.134 | 0.035 | –3.841 | <0.001 | [–0.203, –0.065] | 0.069 | |
| 0.7 | –0.157 | 0.039 | –4.020 | <0.001 | [–0.233, –0.080] | 0.075 | |
| 0.8 | –0.104 | 0.043 | –2.436 | 0.016 | [–0.188, –0.020] | 0.029 | |
| 0.9 | –0.106 | 0.050 | –2.124 | 0.035 | [–0.205, –0.008] | 0.022 | |
| 1.0 | –0.103 | 0.044 | –2.345 | 0.020 | [–0.190, –0.016] | 0.027 |