Evaluating the applicability of replication success metrics in animal-to-human translation: A simulation study

  1. Carolyne Jie Huang
  2. Samuel Pawel
  3. Kimberley Elaine Wever
  4. Benjamin Victor Ineichen
  5. Rachel Heyard  Is a corresponding author
  1. Master Program in Biostatistics, Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Switzerland
  2. Newborn Research, Department of Neonatology, University and University Hospital Zurich, Switzerland
  3. Department of Biostatistics, Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Switzerland
  4. Center for Reproducible Science and Research Synthesis, Epidemiology, Biostatistics and Prevention Institute, University of Zurich, Switzerland
  5. Department of Anesthesiology, Pain and Palliative Medicine, Radboud University Medical Center, Netherlands
  6. Department of Clinical Research, University of Bern, Switzerland
5 figures, 3 tables and 1 additional file

Figures

Proportion of statistically significant findings and average animal effect size.

(a) Nested loop plot of the proportion of statistically significant animal and human findings over all simulation repetitions depending on the simulation conditions, that is animal and human effect sizes, heterogeneity across animal and human studies, animal study sample sizes, and the number of animal studies pooled together to obtain the animal finding. The dotted horizontal lines represent a proportion of 2.5% and 80%. The legend under each plot shows which of the progressively thinner columns in the plot correspond to which combination of simulation conditions. Each horizontal line segment contains the proportion of significant findings under each combination of conditions. For example, the segment highlighted with the arrow represents the proportion of significant findings in the animal studies when the smaller sample size was used, the animal effect size was small, there was no heterogeneity across the animal studies, and five animal studies were pooled. (b) Nested loop plot of the average animal effect size over all simulation repetitions, depending on the simulation conditions and the decision criterion applied to the animal finding. The average effect size observed in the human studies is not affected by the applied criterion. Note that since the criterion was not added as a simulation condition, the represented data is correlated, as the same simulation repetitions are used to calculate the average effect size for the strict, lenient, and no criterion.

Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.

Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. Note that the results for the replication BF and the meta-analysis are not shown here for better readability. The dotted horizontal lines represent α2=0.000625, α=0.025, 1β=0.8, and (1β)2=0.64. All animal studies in this representation are simulated with a small sample size per group (nA=10).

Appendix 1—figure 1
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.

Each of the plots in the grid represent another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. The dotted horizontal lines represent α2=0.000625, α=0.025, 1β=0.8, and (1β)2=0.64. All animal studies in this representation are simulated with a small sample size per group (nA=10).

Appendix 1—figure 2
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.

Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. The dotted horizontal lines represent α2=0.000625, α=0.025, 1β=0.8, and (1β)2=0.64. All animal studies in this representation are simulated with a larger sample size per group ((1β)2=0.64).

Appendix 1—figure 3
Grid of nested loop plots of the proportions of animal-human pairs for which the different metrics flagged successful translation across simulation conditions under no criterion.

Each of the plots in the grid represents another animal-human finding combination. In the first column, for example, the human studies are all simulated under the null hypothesis of no effect. Note that the results for the replication BF and the meta-analysis are not shown here for better readability. The dotted horizontal lines represent α2=0.000625, α=0.025, 1β=0.8, and (1β)2=0.64. All animal studies in this representation are simulated with a larger sample size per group (nA=20).

Tables

Table 1
Summary of the simulation factors used to generate animal and human studies.

These were varied in a fully factorial manner.

ParameterValueDescription
μA{0, −4.44, −24.37}True animal effect size
μH{0, −4.44, −24.37}True human effect size
τA2{0, 19.5, 291.1}Heterogeneity across animal studies
τH2{0, 19.5, 291.1}Heterogeneity across human studies
nA{10, 20}Group sample size for animal studies (small or larger)
k{2, 3, 4, 5}Number of animal studies to pool together
Table 2
Definition and rationale for the implemented decision criteria to continue.
NameDescriptionReal-world analogue
Strict criterionWe progressed to a human study if the random-effects meta-analysis of the corresponding k animal studies found a significant beneficial treatment effect with one-sided significance level α=0.025.Related to regulatory-style evidence thresholds and presents a highly selective preclinical pipeline.
Lenient criterionWe progressed to a human study whenever the random-effects meta-analysis of the k animal studies found a beneficial effect (i.e. if the estimated effect was negative), even if it was not statistically significant at α=0.025.Represents an exploratory decision-making where efficacy evidence is considered in addition to other factors (e.g. safety), but not intended to represent a certain regulatory standard.
No criterionAs a point of reference, we also computed performance measures for all simulated animal and human findings, regardless of the results of the animal studies.Represents a reference scenario mainly used to evaluate metric behavior independent of progression decisions.
Table 3
Summary of simulation results across nine translation success metrics (no criterion).

At baseline refers to no heterogeneity across animal and none across human studies. The proportion P refers to the proportion of pairs of animal and human findings for which the metric flagged translation success (under specific simulation conditions).

MetricStrengthsWeaknessesBehavior with increasing heterogeneitySensitivity to more animal data (larger k and nA)Behavior under effect mismatch (animal vs human)
Significance criterionExcellent control of overall T1E rate at α2 (at baseline); one of the most reliable metrics under heterogeneity in human studies.Lowest translation power in many scenarios; often too conservative when animal findings are low-powered.Tends to be the metric least affected compared to baseline, unless human heterogeneity is high.Larger nA increases P slightly; larger k decreases false translation success rate and increases translation power, unless there is high animal heterogeneity.Independent; it requires animal and human study to be significant; a large animal effect cannot ‘mask’ a null human effect.
Meta-analysisHighest translation power in many conditions; outperforms other metrics when true effects exist in both species.Very high overall T1E rates (up to 40%); flags success even if only one species shows a strong result.Highly sensitive to human heterogeneity (high heterogeneity in human causes a massive increase in false positive translations); less sensitive to animal heterogeneity.Larger nA and k increase P if animal effect is non-null, and decrease it otherwise.Treats animal and human as interchangeable and findings as equivalent, allowing a very convincing animal finding to hide a human null result (and vice-versa).
Replication BFUseful in case of high heterogeneity scenarios (where it becomes more comparable to other metrics).Very low power under no to low heterogeneity scenarios; potentially weights human data too heavily.Very sensitive to changes in heterogeneity; P generally increases with heterogeneity, unless animal effect is null or small and human is small.Larger nA decreases P; an increase in k increases P when true animal and human effect are the same, and decreases it if there is a mismatch, unless there is high heterogeneity in animal study (which inverses the trend).Human-focused; functions a bit like a test of the human study, largely ignoring the original animal evidence.
Edgington (unweighted / weighted)Better translation power than significance criterion while maintaining good overall T1E rate control at baseline; P for weighted version pulled up or down depending on evidence in human studies.Generally conservative; weighted version can have slightly higher overall T1E rates than significance criterion.Generally follows significance criterion with slightly higher proportions; unweighted and weighted come closer together with high animal heterogeneity.Larger nA generally increases proportions unless there is no heterogeneity and a null effect in the animal studies; proportions increase slightly with k if animal effect is non-zero.Additive; requires evidence from both but allows a very strong result in one to compensate for a slightly weaker result in the other.
Golden skeptical p-value (No shrinkage)Best at keeping overall T1E rates low when animal study heterogeneity is high.Lowest translation power among skeptical p metrics and lowest translation power in small effect scenarios; penalizes any discrepancy in effect size.Sensitive, especially to high heterogeneity across animal studies.Translation power increases with k due to higher chance of a significant animal finding unless there is high animal heterogeneity; P increases with nA.Penalizing; specifically designed to flag failure if the human effect is smaller than the animal effect (shrinkage) and one of the studies is not convincing enough.
Golden skeptical p-value (Mod/High shrink)High translation power; allows weaker animal findings to translate when human findings are strong and vice-versa, even if shrinkage is present.Higher overall T1E rates; less conservative to shrinkage (by design).Sensitive to human heterogeneity; similar trend compared to no shrinkage version.Highly sensitive; translation power increases significantly with k as it decreases the relative sample size ratio; P increases with nA.Adaptive; permits for potentially realistic shrinkage between animal and human effects without automatically flagging failure.
Controlled skeptical p-valueMost consistently high translation power across varying conditions.Higher overall T1E rate than significance criterion; similar overall T1E patterns to high-shrinkage metrics.Moderate sensitivity; follows the overall T1E rate of human studies as heterogeneity increases.Strong response to k; increasing k to 5 allows the metric to reach acceptable translation power even with noisy animal data; P increases with nA.Balanced; allows for effect size differences but requires a convincing level of evidence that accounts for the sample size ratio.

Additional files

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Carolyne Jie Huang
  2. Samuel Pawel
  3. Kimberley Elaine Wever
  4. Benjamin Victor Ineichen
  5. Rachel Heyard
(2026)
Evaluating the applicability of replication success metrics in animal-to-human translation: A simulation study
eLife 15:RP109853.
https://doi.org/10.7554/eLife.109853.3