Evaluating the applicability of replication success metrics in animal-to-human translation: A simulation study
Figures
Proportion of statistically significant findings and average animal effect size.
(a) Nested loop plot of the proportion of statistically significant animal and human findings over all simulation repetitions depending on the simulation conditions, that is animal and human effect sizes, heterogeneity across animal and human studies, animal study sample sizes, and the number of animal studies pooled together to obtain the animal finding. The dotted horizontal lines represent a proportion of 2.5% and 80%. The legend under each plot shows which of the progressively thinner columns in the plot correspond to which combination of simulation conditions. Each horizontal line segment contains the proportion of significant findings under each combination of conditions. For example, the segment highlighted with the arrow represents the proportion of significant findings in the animal studies when the smaller sample size was used, the animal effect size was small, there was no heterogeneity across the animal studies, and five animal studies were pooled. (b) Nested loop plot of the average animal effect size over all simulation repetitions, depending on the simulation conditions and the decision criterion applied to the animal finding. The average effect size observed in the human studies is not affected by the applied criterion. Note that since the criterion was not added as a simulation condition, the represented data is correlated, as the same simulation repetitions are used to calculate the average effect size for the strict, lenient, and no criterion.
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.
Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. Note that the results for the replication BF and the meta-analysis are not shown here for better readability. The dotted horizontal lines represent , , , and . All animal studies in this representation are simulated with a small sample size per group ().
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.
Each of the plots in the grid represent another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. The dotted horizontal lines represent , , , and . All animal studies in this representation are simulated with a small sample size per group ().
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.
Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. The dotted horizontal lines represent , , , and . All animal studies in this representation are simulated with a larger sample size per group ().
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.
Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. Note that the results for the replication BF and the meta-analysis are not shown here for better readability. The dotted horizontal lines represent , , , and . All animal studies in this representation are simulated with a larger sample size per group ().
Tables
Summary of the simulation factors used to generate animal and human studies.
These were varied in a fully factorial manner.
| Parameter | Value | Description |
|---|---|---|
| {0, −4.44, −24.37} | True animal effect size | |
| {0, −4.44, −24.37} | True human effect size | |
| {0, 19.5, 291.1} | Heterogeneity across animal studies | |
| {0, 19.5, 291.1} | Heterogeneity across human studies | |
| {10, 20} | Group sample size for animal studies (small or larger) | |
| {2, 3, 4, 5} | Number of animal studies to pool together |
Definition and rationale for the implemented decision criteria to continue.
| Name | Description | Real-world analogue |
|---|---|---|
| Strict criterion | We progressed to a human study if the random-effects meta-analysis of the corresponding animal studies found a significant beneficial treatment effect with one-sided significance level . | Related to regulatory-style evidence thresholds and presents a highly selective preclinical pipeline. |
| Lenient criterion | We progressed to a human study whenever the random-effects meta-analysis of the animal studies found a beneficial effect (i.e. if the estimated effect was negative), even if it was not statistically significant at . | Represents an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g. safety), but not intended to represent a certain regulatory standard. |
| No criterion | As a point of reference, we also computed performance measures for all simulated animal and human findings, regardless of the results of the animal studies. | Represents a reference scenario mainly used to evaluate metric behavior independent of progression decisions. |
Summary of simulation results across nine translation success metrics (no criterion).
At baseline refers to no heterogeneity across animal and none across human studies. The proportion P refers to the proportion of pairs of animal and human findings for which the metric flagged translation success (under specific simulation conditions).
| Metric | Strengths | Weaknesses | Behavior with increasing heterogeneity | Sensitivity to more animal data (larger and ) | Behavior under effect mismatch (animal vs human) |
|---|---|---|---|---|---|
| Significance criterion | Excellent control of overall T1E rate at (at baseline); one of the most reliable metrics under heterogeneity in human studies. | Lowest translation power in many scenarios; often too conservative when animal findings are low-powered. | Tends to be the metric least affected compared to baseline, unless human heterogeneity is high. | Larger increases P slightly; larger decreases false translation success rate and increases translation power, unless there is high animal heterogeneity. | Independent; it requires animal and human study to be significant; a large animal effect cannot ‘mask’ a null human effect. |
| Meta-analysis | Highest translation power in many conditions; outperforms other metrics when true effects exist in both species. | Very high overall T1E rates (up to 40%); flags success even if only one species shows a strong result. | Highly sensitive to human heterogeneity (high heterogeneity in human causes a massive increase in false positive translations); less sensitive to animal heterogeneity. | Larger and increase P if animal effect is non-null, and decrease it otherwise. | Treats animal and human as interchangeable and findings as equivalent, allowing a very convincing animal finding to hide a human null result (and vice-versa). |
| Replication BF | Useful in case of high heterogeneity scenarios (where it becomes more comparable to other metrics). | Very low power under no to low heterogeneity scenarios; potentially weights human data too heavily. | Very sensitive to changes in heterogeneity; P generally increases with heterogeneity, unless animal effect is null or small and human is small. | Larger decreases P; an increase in increases P when true animal and human effect are the same, and decreases it if there is a mismatch, unless there is high heterogeneity in animal study (which inverses the trend). | Human-focused; functions a bit like a test of the human study, largely ignoring the original animal evidence. |
| Edgington (unweighted / weighted) | Better translation power than significance criterion while maintaining good overall T1E rate control at baseline; P for weighted version pulled up or down depending on evidence in human studies. | Generally conservative; weighted version can have slightly higher overall T1E rates than significance criterion. | Generally follows significance criterion with slightly higher proportions; unweighted and weighted come closer together with high animal heterogeneity. | Larger generally increases proportions unless there is no heterogeneity and a null effect in the animal studies; proportions increase slightly with if animal effect is non-zero. | Additive; requires evidence from both but allows a very strong result in one to compensate for a slightly weaker result in the other. |
| Golden skeptical p-value (No shrinkage) | Best at keeping overall T1E rates low when animal study heterogeneity is high. | Lowest translation power among skeptical metrics and lowest translation power in small effect scenarios; penalizes any discrepancy in effect size. | Sensitive, especially to high heterogeneity across animal studies. | Translation power increases with due to higher chance of a significant animal finding unless there is high animal heterogeneity; P increases with . | Penalizing; specifically designed to flag failure if the human effect is smaller than the animal effect (shrinkage) and one of the studies is not convincing enough. |
| Golden skeptical p-value (Mod/High shrink) | High translation power; allows weaker animal findings to translate when human findings are strong and vice-versa, even if shrinkage is present. | Higher overall T1E rates; less conservative to shrinkage (by design). | Sensitive to human heterogeneity; similar trend compared to no shrinkage version. | Highly sensitive; translation power increases significantly with as it decreases the relative sample size ratio; P increases with . | Adaptive; permits for potentially realistic shrinkage between animal and human effects without automatically flagging failure. |
| Controlled skeptical p-value | Most consistently high translation power across varying conditions. | Higher overall T1E rate than significance criterion; similar overall T1E patterns to high-shrinkage metrics. | Moderate sensitivity; follows the overall T1E rate of human studies as heterogeneity increases. | Strong response to ; increasing to 5 allows the metric to reach acceptable translation power even with noisy animal data; P increases with . | Balanced; allows for effect size differences but requires a convincing level of evidence that accounts for the sample size ratio. |