Figures and data

The permutation test, perfect match concept, and permute-match tests.
(A) The permutation test. Circles indicate the average correlation without trial swapping (ρorig; pink) or with trial swapping (ρshuffle; grey). Here, N≥ is the number of average correlations (including ρorig) that are equal to or larger than the original correlation ρorig. It follows that pperm = N≥/n! because n! is the number of ways the trials can be shuffled. (B) Illustration of a Y-perfect match, represented as a table of correlations wherein each diagonal entry is the greatest in its own row. Each line with a “>” or “<” symbol denotes a required ordering relationship between the two numbers at either end of the line. Note that this example has a Y-perfect match, but not an X-perfect match. For an X-perfect match, each diagonal entry would need to be the greatest in its own column. (C) If an X- or Y-perfect match occurs, then the original correlation ρorig is always greater than any shuffled correlation ρshuffle. Top: If a Y-perfect match occurs, then each Yi gives the greatest correlation when it is paired with the Xi from the same trial. A shuffled correlation ρshuffle pairs Yis from some trials with Xjs from different trials, thereby reducing the correlation. Lines with symbols “>”, “<”, and “=” denote comparisons between terms. Bottom: If an X-perfect match occurs, a similar argument can be applied. (D) Since a perfect match ensures that the original correlation is greater than any shuffled correlation, a perfect match also ensures that pperm takes on its lowest possible value of 1/n!. (E) The simultaneous permute-match(X;Y) test. (F) The sequential permute-match(X▶Y) test; note that the permute-match(Y ▶X) test is defined analogously.

Statistical power of various permutation test variants as illustrated using a simple nonstationary system.
(A) Summary of test power as a function of the significance level α and the number of replicates n. (B, C) System equations and example dynamics. The processes X and Y are given by a linear trend with additive noise on a time grid t = 1, 2, …, 100. The noise terms ϵX,i(t) and ϵY,i(t) are drawn from a bivariate normal distribution with a mean of 0, variance of 1, and covariance of rX,Y. The chart shows an example pair of time series where X and Y are dependent (rX,Y = 0.3). (D) Statistical power for the permutation test and various permute-match tests as a function of the replicate number n, significance level α, and strength of dependence rX,Y. Power was estimated as the proportion of simulations in which dependence was detected, calculated from 5000 simulations at each value of rX,Y between rX,Y = 0 and 0.54 in steps of size 0.01. At rX,Y = 0, there is no dependence, so the curve at that point indicates the false positive rate rather than power. We chose the Pearson correlation coefficient as our correlation function ρ. See (A) for the color legend.

A sequential permute-match test detects dependence between average individual speed and alignment (as measured by mean resultant length) in small groups of zebrafish.
(A) Time series of average individual speed and mean resultant length for the three replicate videos. Black curves show original time series, and red curves show a 30-second moving average. Average individual speed appears to decrease with time in all trials. All 12 time series (2 variables × 3 trials × 2 smoothing conditions) were deemed non-stationary by a Kwiatkowski-Phillips-Schmidt-Shin (KPSS; [37]) test (p < 0.01). Note that the KPSS test seeks to reject a stationary null hypothesis in contrast to other common tests, whose null hypotheses are nonstationary. (B) Pearson correlation between alignment and average individual speed, both within trials and between trials, using all frames up to the length of the shortest speed and alignment time series (19308 frames; 10.06 minutes). Table entries are shaded by correlation. We observe a speed-perfect match (and an alignment-perfect match), and thus detect dependence with p = 1/33 ≈ 0.04. (C) A parametric test also detects dependence. As a parametric alternative, for each trial we averaged values of speed and mean resultant length over the complete time series (visualized in the scatter plot shown, where each point is one trial). The sample Pearson correlation of the time-averaged variables is 0.99996, and a one-tailed test of significance gives p ≈ 0.003. (D) Permute-match tests detected a significant correlation between speed and alignment more consistently than the parametric test. For a grid of lengths between 20 and 600 seconds we sampled 500 random segments of each length, each drawn from the first 600 seconds, and determined for each segment whether the parametric test and/or the two possible permute-match tests detected a significant (p ≤ 0.05) correlation. In the edge case of the maximum 600-second length, all 500 “random” segments were identical. We used the Pearson correlation as the correlation function ρ in the permute-match tests, and a one-tailed parametric test, reflecting the alternative hypothesis of correlation being positive (not merely nonzero).

The mean resultant length is a measure of alignment among angles.
Consider three unit vectors colored blue, yellow, and red. The length of net vector (dashed) divided by the number of vectors is defined as the mean resultant length. Compared to less-aligned angles (i), strongly aligned angles (ii) produce a longer net vector, and thus a greater mean resultant length.

Joint and marginal probabilities of the outcomes of tests X and Y, assuming that pX and pY both follow Unif(0, 1).
We first fill the total probability row and column, and define P′ as the probability that test Y is significant but test X is not. The remaining entries can then be deduced.

The procedure of Eq S15 produces a false positive rate that generally exceeds α, and depends on the relationship between tests X and Y.
This relationship between X and Y can be quantified by P′, the probability that test Y is significant (i.e. 2pY ≤ α) but test X is not significant (i.e. pX >α). The false positive rate is shown over the full range of P′, from 0 to α/2. The only way for the overall procedure of Eq S15 to be valid is when P′ = 0, where tests X and Y are so tightly coupled that a significant result in test Y deterministically ensures that test X is also significant. Other than this edge case, the false positive rate of the overall procedure will exceed α.

Statistical power of permute-match tests and the permutation test in a nonlinear and nonstationary system.
(A) System equations. (B) Example dynamics. The processes X and Y are given by a coupled logistic map on a time grid t = 1, 2, …, 100. The system is nonstationary because the parameter r(t) varies with time from 3.72 when t = 1 to 3.82 when t = 99. (C) To see the nonstationarity clearly, we plot Xi(t) against Xi(t + 1) and color points by time, showing that the parabola that the points lie on is drifting over time. To better show the trend, 10 replicates are shown simultaneously in this chart. (D) Statistical power of the permutation test and permute-match tests as a function of the number of replicates, the significance level, and rX,Y. Power was estimated from 5000 simulations at each value of rX,Y between rX,Y = 0 and rX,Y = 0.2 in steps of size 0.005. The correlation statistic (ρ) was cross-map skill, which is known to readily detect dependence between X and Y in this system [46]. For the cross-map skill calculation, we used Y to estimate X, which corresponds to a scenario where the data analyst hypothesizes that X influences Y. Cross-map skill requires two parameters, the embedding dimension and the embedding lag, and these were set to 2 and 1 respectively following prior works that used the logistic map for benchmarking [46, 47].
