Figures and data

Quantification of individual prediction tendency and the multi-speaker paradigm.
random). Four tones of different fundamental frequencies were presented with a fixed stimulation rate of 3 Hz, their transitional probabilities varied according to respective conditions. B) Expected classifier decision values contrasting the brains’ prestimulus tendency to predict a forward transition (ordered vs. random). The purple shaded area represents values that were considered as prediction tendency C) Exemplary excerpt of a tone sequence in the ordered condition. An LDA classifier was trained on forward transition trials of the ordered condition (75% probability) and tested on all repetition trials to decode sound frequency from brain activity across time. D) Participants either attended to a story in clear speech, i.e. 0 distractor condition, or to a target speaker with a simultaneously presented distractor (blue), i.e. 1 distractor condition. E) The speech envelope was used to estimate neural and ocular speech tracking in respective conditions with temporal response functions (TRF). F) The last noun of some sentences was replaced randomly with an improbable candidate to measure the effect of envelope encoding on the processing of semantic violations. Adapted from Schubert et al., 2023.

List of all ebooks and short stories that were used as a basis for audio material.
Note: each story was narrated by a target speaker (2 x ∼3 min excerpts each), and a distractor speaker (1 x ∼3 min different excerpt of the same story).

Individual prediction tendency.
A) Time-resolved contrasted classifier decision: forward > repetition for ordered and random repetition trials. Classifier tendencies showing frequency-specific prediction for tones with the highest probability (forward transitions) can be found even before stimulus onset but only in an ordered context (shaded areas always indicate 95% confidence intervals). Using the summed difference across pre-stimulus time, one prediction value was extracted per individual subject. B) Distribution of prediction tendency values across subjects (N = 29).

Neural speech tracking is related to prediction tendency and word surprisal, independent of selective attention.
The TRF (filter kernel, h) models how the brain processes the envelope over time. This filter is used to predict neural responses via convolution. Predicted responses are correlated with actual neural activity to evaluate model fit and the TRF’s ability to capture response dynamics. Correlation coefficients from these models are then used as dependent variables in Bayesian regression models. (Panel adapted from Gehmacher et al., 2024b). B) Temporal response functions (TRFs) depict the time-resolved neural tracking of the speech envelope for the single speaker and multi speaker target condition, shown here as absolute values averaged across channels. Solid lines represent the group average. Shaded areas represent 95% Confidence Intervals. C–H) The beta weights shown in the sensor plots are derived from Bayesian regression models in A. For Panel C, this statistical model is based on correlation coefficients computed from the TRF models (further details can be found in the Methods Section). C) In a single speaker condition, neural tracking of the speech envelope was significant for widespread areas, most pronounced over auditory processing regions. D) The condition effect indicates a decrease in neural speech tracking with increasing noise (1 distractor). E) Stronger prediction tendency was associated with increased neural speech tracking over left frontal areas. F) However, there was no interaction between prediction tendency and conditions of selective attention. G) Increased neural tracking of semantic violations was observed over left temporal areas. H) There was no interaction between word surprisal and speaker condition, suggesting a representation of surprising words independent of background noise. Marked sensors indicate ‘significant’ clusters, defined as at least two neighboring channels showing a significant result. N = 29.

Ocular speech tracking is dependent on selective attention.
Temporal profiles of this effect show a downward pattern (negative TRF weights). B) Horizontal eye movements ‘significantly’ track attended speech in a multi-speaker condition. Temporal profiles of this effect show a left-rightwards (negative to positive TRF weights) pattern. Statistics were performed using Bayesian regression models. A ‘*’ within posterior distributions depicts a significant difference from zero (i.e. the 94%HDI does not include zero). Shaded areas in TRF weights represent 95% confidence intervals. N= 29

Model summary statistics for ocular speech tracking depending on condition and prediction tendency.
Note: Dependent Variable = ocular speech tracking (correlation between true and predicted eye movements).

Model summary statistics for ocular speech tracking depending on word type and condition.
Note: Dependent Variable = ocular speech tracking (correlation between true and predicted eye movements * Intercept represents speech tracking for lexical control words in a single speaker condition.).

Ocular speech tracking and selective attention to speech share underlying neural computations.
(A) Vertical eye movements mediate neural tracking of clear speech. Top: Spatial distribution (PCA weights) of the group-averaged mediation effect for the first three principal components (PC1–PC3). Mediation effects are distributed across bilateral language networks, including early sensory, parietal, frontal, and lateral semantic processing areas. Middle: Amplitude of the Total Temporal Response Function (TRF; green) versus the Direct TRF (grey) over time. Bottom: Time-resolved standardized beta coefficients representing the mediation effect. Vertical eye movements significantly mediate neural tracking at early latencies (0.07–0.28 s). (B) Horizontal eye movements contribute to the neural tracking of target speech in a multi-speaker environment. Layout is identical to (A), with Total TRF shown in blue. Horizontal eye movements show similar early mediation effects alongside two additional negative clusters, indicating later effects at ∼300 ms and ∼400 ms respectively. Shaded ribbons around the waveform lines indicate mean and standard error for the TRF comparison panels and the mean and 94% HDI for the mediation effect panels. In the bottom panels, the grey shaded rectangular background band represents the Region of Practical Equivalence (ROPE; Kruschke, 2018). Solid horizontal grey markers denote robust clusters where a minimum of two contiguous time-points exhibited a significant mediation effect. Statistical evaluations were performed using Bayesian regression models (N = 29).

Ocular, but not neural speech tracking is related to semantic speech comprehension.
B & C) A ‘significant’ negative relationship between comprehension and vertical as well as horizontal ocular speech tracking shows that participants with weaker comprehension increasingly engaged in ocular speech tracking. Statistics were performed using Bayesian regression models. Shaded areas represent 94% HDIs. N = 29

Model summary statistics for comprehension depending on ocular speech tracking and condition.
Note: Dependent Variable = comprehension (averaged accuracy for comprehension questions). * Intercept represents comprehension for average speech tracking in a single speaker condition.

Model summary statistics for rated difficulty depending on ocular speech tracking and condition.
Note: Dependent Variable = rated difficulty (averaged subjective ratings on a 5-point likert scale). * Intercept represents rated difficulty for average speech tracking in a single speaker condition.

A schematic illustration of the framework.
The temporal and spatial characteristics of a representation depends on where and when (in the brain) it is probed. Anticipatory predictions (purple) help to interpret auditory information at different levels in parallel with high feature-specificity but low temporal precision. These anticipatory predictions reflect to some extent individual tendencies (and differences) that generalise across listening situations. In contrast, active ocular sensing (green) increases the temporal precision already at lower stages of the auditory system to facilitate bottom–up processing at specific timescales (similar to neural oscillations). It does not necessarily convey feature-specific information, but is more likely used to boost (or filter) information around relevant time windows. Our results suggest that this mechanism is motivated by selective attention (blue) rather than predictive assumptions