Sensitivity of the human temporal voice areas to nonhuman primate vocalizations

eLife Assessment

This important study shows that regions of the human auditory cortex that respond strongly to human voices are also sensitive to vocalizations from closely related primate species. The evidence is convincing and methodologically strong. The work offers significant insight into the evolutionary continuity of voice processing and would be of interest to researchers studying auditory processing and evolutionary neuroscience in general.

https://doi.org/10.7554/eLife.108795.3.sa0

Abstract

In recent years, research on voice processing in the human brain, particularly the study of temporal voice areas (TVAs), was dedicated almost exclusively to conspecific vocalizations. To characterize commonalities and differences regarding primate vocalization representations in the human brain, the inclusion of closely related nonhuman primates, namely chimpanzees and bonobos, is needed. We hypothesized that neural commonalities would depend on both phylogenetic and acoustic proximities, with chimpanzees ranking closest to Homo. Presenting human participants (N = 23) with the vocalizations of four primate species (rhesus macaques, chimpanzees, bonobos, and humans) and regressing-out relevant acoustic parameters using three distinct analyses, we observed within-TVA, sample-specific, bilateral anterior superior temporal gyrus activity for chimpanzee vocalizations compared to: all other species; nonhuman primates; and human vocalizations. Within-TVA activity was also observed for macaque vocalizations. Our results provide evidence for subregions of the TVA that respond principally, but not exclusively, to phylogenetically and acoustically close nonhuman primate vocalizations, namely those of chimpanzees.

Introduction

The study of cerebral mechanisms underlying speech and voice processing has gained importance since the early 2000s with the advent of functional magnetic resonance imaging (fMRI) (Ogawa et al., 1990). Voice-sensitive areas, commonly referred to as ‘temporal voice areas’ (TVAs) or simply ‘voice areas’, have been highlighted along the upper, superior portion of the temporal cortex (Belin et al., 2000). Since then, great efforts have been made to better characterize these TVA, with particular attention to their spatial division into functional subregions (Aglieri et al., 2018; Kriegstein and Giraud, 2004; Pernet et al., 2015). A fairly large body of literature points to the critical role of the TVA in voice perception and processing in healthy participants (Kriegstein and Giraud, 2004; Frühholz et al., 2020; Latinus and Belin, 2011; Zäske et al., 2017) as well as in lesioned patients (Roswandowitz et al., 2018). Subregions of the TVA have also been directly linked to social perception (Lahnakoski et al., 2012), vocal emotion processing (Ethofer et al., 2012; Witteman et al., 2012), voice identity (Latinus et al., 2011; Latinus et al., 2013), and gender perception (Charest et al., 2013). The developmental axis of voice processing has also been studied in infants, demonstrating the existence of the TVA and voice sensitivity in the human brain as early as 4 (Calce et al., 2024) and 7 (Grossmann et al., 2010) months of age, while the ability to respond specifically to the voice of their parents has been observed in fetuses in utero (Kisilevsky et al., 2003). With the ongoing development of brain imaging and analysis techniques (Hüppi, 2011), it is realistic to expect successful, albeit noninvasive, fMRI results on task-related voice perception in utero in the near future. Along the evolutionary axis, evidence for TVA or, more generally, conspecific vocalization-sensitive brain areas has emerged primarily in dogs (Andics et al., 2014), monkeys (Perrodin et al., 2011; Petkov et al., 2008) (Macaca mulatta) as well as recently in common marmosets (Dureux et al., 2025; Jafari et al., 2023) (Callithrix jacchus), raising the question of whether such specialized brain areas are species-specific (Fecteau et al., 2004) and to what extent human and nonhuman primates share neural mechanisms that enable them to preferentially process conspecific vocalizations (Belin, 2006). However, less attention has been paid to paradigms in which animal vocalizations are presented to humans—see Fecteau et al., 2004's study on species-specificity in the human brain. Human processing of animal vocalizations has been studied with both monkey and cat material, but no specific cross-species activations have been observed within the TVA with respect to either species (Belin et al., 2008a). Other studies have focused more specifically on phylogenetic distance and have included nonhuman ape (chimpanzee, Pan troglodytes) and ‘Old World’ monkey (rhesus macaque, M. mulatta) vocalizations as stimuli. Such studies failed to identify species-specific brain activations—despite correctly discriminating chimpanzee affective vocalizations (Fritz et al., 2018)—and observed ambivalent results for below (Fritz et al., 2018) vs. above (Linnankoski et al., 1994) chance discrimination of macaque affective vocalizations by human participants. A recent exception is a study in which functionally homologous anterior TVA activity was observed in both humans and macaques: this region was indeed specific to macaque calls in the macaque’s anterior TVA, and specific to human voices in the anterior TVA of humans, but no macaque-specific activity was observed in the human TVA (Bodin et al., 2021). Additionally, TVA responded similarly in both species to the presentation of nonverbal auditory stimuli, potentially due to some level of acoustic similarity (Bodin et al., 2021). This sparse literature motivated the present study, which aims to investigate cross-species TVA activations in humans when asked to categorize vocalizations from phylogenetically—and acoustically close and distant—species while undergoing fMRI scanning. While studies on the TVA typically involve passive listening, we wanted to additionally have the means to compare nonhuman primate species recognition/categorization in humans. The importance of acoustic differences between species and more specifically acoustic distance, particularly through fundamental frequency variations (Grawunder et al., 2018; Slocombe and Zuberbühler, 2007), was indeed of great interest. Acoustic distance—calculated using Mahalanobis distance with 16 acoustic parameters extracted from the stimuli, see Supplementary file 1A—was in fact a determining parameter in assessing affective cues recognition in nonhuman primate calls by human participants (Debracque et al., 2023). In this study, affiliative chimpanzee—but not bonobo—calls were acoustically the closest to positive human voice stimuli, suggesting a distinct evolution of bonobo calls (Debracque et al., 2023). Bonobo vocalizations are of particular interest because this species is thought to have undergone evolutionary changes in their communication, in part due to a neoteny process involving acoustic modifications, and although they are as phylogenetically close to humans as chimpanzees—with an estimated separation with the Homo lineage only 6–8 million years ago (Gruber and Clay, 2016). Previous research has shown that bonobos have a shorter larynx—a valid predictor of a species’ mean fundamental frequency (Titze et al., 2016)—compared to chimpanzees, resulting in a higher fundamental frequency in their calls (Grawunder et al., 2018). Such a difference has been demonstrated in juvenile bonobo calls compared to chimpanzee and human baby calls (Kelly et al., 2017), arguing for a greater acoustic distance between bonobo calls and human or chimpanzee vocalizations. For these reasons, we included vocalizations from both Pan species (chimpanzees, P. troglodytes; bonobos, Pan paniscus), as well as a phylogenetically more distant species (Cercopithecidae: rhesus monkeys), with an estimated separation with the Homo lineage dating back to 25 million years ago. Indeed, any claim of human ‘uniqueness’ for TVA selectivity remains on hold and should be tested in light of these closely related species. Using the same stimuli, we previously investigated the frontal mechanisms involved in the categorization of nonhuman primate vocalizations independently of a selection of low-level acoustic parameters (Ceravolo et al., 2023), but the possibility that acoustic differences would affect, at the auditory level, the ability of human participants to recognize nonhuman primate calls should be thoroughly examined, as we did in the present study. As suggested by research mentioned above, monkey vocalizations are overall less likely to be identified compared to ape vocalizations due to both phylogenetic and acoustic differences. Therefore, our mechanistic hypothesis of the difficulty for humans to recognize bonobo calls is that frequencies of the human tonotopic map in the auditory cortex—adapted and adjusted to the frequencies of the human voice during evolution—would not be tailored to process the frequencies generated by bonobo calls. It would also be the case for macaque calls, while frequencies of chimpanzee calls—being closer to the range of human voice fundamental frequency (Grawunder et al., 2018; Debracque et al., 2023)—would be better represented in the human auditory cortex and therefore more easily processed and better identified by humans.

According to the literature mentioned so far and to the mechanistic hypothesis underlying the processing of chimpanzee as opposed to bonobo and/or macaque calls by human participants, we therefore predicted: (1) Bioacoustics: more acoustic proximity between human and chimpanzee vocalizations, whereas more distance would separate those of bonobos and macaques from the human voice (Grawunder et al., 2018; Debracque et al., 2023; Cauzinille et al., 2024); (2) Phylogeny: a selectivity of the superior temporal lobe—within the TVA—for the processing of vocalizations from the Pan taxon (chimpanzee, bonobo) but not Cercopithecidae (rhesus monkey) vocalizations (Fecteau et al., 2004; Fritz et al., 2018; Gruber and Clay, 2016; Kelly et al., 2017; Debracque et al., 2025)—while considering acoustic features that characterize the stimuli the best through a discriminant analysis.

Results

Behavior: general aspects

Our hypotheses involve a control of phylogeny through the inclusion of specific primate species calls as well as the selection of specific acoustic features. We programmed a task in which the vocalizations of each species were presented randomly, and for which the participants (N = 23) had to specify to which species each stimulus corresponded. We therefore included equal numbers of trials (N = 72) with human, chimpanzee, bonobo, and macaque vocalizations—N = 18 each—as well as trial-level acoustic features of the vocalizations, using distinct statistical models with specific covariates. The stimuli were presented at a constant sound pressure level, but the stimuli were not normalized to avoid altering their naturality (Ferdenzi et al., 2013). The statistical models are sorted from the least to the most sophisticated modeling to uncover the role(s) of acoustic features on TVA activity potentially specific to each/some species (see the Methods; Model 1: mean of vocalization fundamental frequency and energy; Model 2: multi-dimensional Mahalanobis acoustic distance between the human voice and the calls of each nonhuman primate species De Maesschalck et al., 2000; Model 3: between-species most discriminant acoustic features of our stimuli, extracted using a general discriminant analysis GDA; Debracque et al., 2023; Model 4: model-based approach of the probability of correct species categorization, including the most discriminant acoustic features of Model 3). Acoustical analyses involved in Model 2 allowed us to validate our first hypothesis according to which chimpanzees are acoustically the closest to humans, followed by the calls of bonobos and macaques (Figure 1C)— the main effect of species on the acoustic distance was significant, F(3,88) = 15.84, p < 0.001, as well as all comparisons (Hum > Chimp: t(23) = −12.22, p < 0.001; Hum > Bon: t(23) = −14.16, p < 0.001; Hum > Mac: t(23) = −22.57, p < 0.001; Hum > (Chimp, Bon, Mac): t(23) = −20.86, p < 0.001; Chimp > Bon: t(23) = −5.07, p < 0.01; Bon > Mac: t(23) = −6.43, p < 0.001; Chimp > (Bon, Mac): t(23) = −27.34, p < 0.001; see also Supplementary file 1B).

Timecourse and results of the species categorization task and the control task testing for potential attentional biases.

(A) Detail of the timecourse of four trials of the species categorization task in non-representative order, including waveform and spectrogram graphs for one example stimulus of each species. (B) Behavioral results of the task (N = 23) showing the probability (Probab.) of correctly categorizing (categ.) each species’ vocalization, using the six acoustic covariates from Model 3 and reaction times as variables. (C) Histograms of the acoustic Mahalanobis distance data of each species including mean (numbers represent exact mean value). (D) Control task and its design in an independent sample of 28 participants using each species’ vocalization as an exogenous cue preceding the target to be detected as fast as possible, namely a short sine wave tone (600 Hz). (E) Results of the control task (N = 28), showing no biasing of attention by any of the species stimuli: no species triggered attentional capture, yielding to no attentional advantage for target detection. For the results plots, violin plots illustrate distribution fit, error bars the standard error of the mean (SEM). Points represent individual values. ITI: intertrial interval; Resp.: response; Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. **p < 0.01, ***p < 0.001, n.s.: non-significant.

Behavioral data: species categorization task

As described above, Model 4 was computed to analyze the ability of our participants to categorize the species vocalizations, considering their production social contexts, namely agonistic (threat, distress) or affiliative (‘positive’). These analyses used as additional predictors the six most discriminant acoustic features described in Model 3, in addition to controlling for response reaction times (Figure 1A, B). These data revealed a main effect of the species and context factors (χ2(3) = 187.35, p < 0.0001 and χ2(2) = 6.62, p < 0.05) as well as their interaction (χ2(6) = 14.02, p < 0.05). Main effects of acoustic features were also significant: F0 power (χ2(1) = 4.44, p < 0.05) and intensity contour difference (χ2(1) = 6.07, p < 0.05). Species categorization results were explained by human voices being categorized more accurately than all other species (human vs nonhuman primate species: χ2(1) = 144.31, p < 0.0001; human vs chimpanzee: χ2(1) = 97.09, p < 0.0001; human vs bonobo: χ2(1) = 177.91, p < 0.0001; human vs macaque: χ2(1) = 98.34, p < 0.0001) and by worse, below chance-level categorization of bonobo calls compared to all other species (bonobo vs human: χ2(1) = 177.91, p < 0.0001; bonobo vs chimpanzee: χ2(1) = 22.86, p < 0.0001; bonobo vs macaque: χ2(1) = 28.84, p < 0.0001). Interestingly, no difference in accuracy was observed for the categorization of chimpanzee compared to macaque calls (χ2(1) = 0.004, p = 0.94). Since context was controlled for but not part of our hypotheses, and since it was strongly interacting with species, we report all post hoc contrasts in Table 1.

Table 1
Post hoc contrasts for behavioral effects of the species by context interaction.
ContrastEstimateSEdfz ratiop-value(unc)p-value(FDR)
1bon ago–chimp ago–1.786446350.4294698Inf–4.159655763.187276e−056.010292e−05
2bon ago–hum ago–6.300502240.7182617Inf–8.771875121.757152e−182.899301e−17
3bon ago–mac ago–2.008753490.4126352Inf–4.868109971.126706e−062.478754e−06
4bon ago–bon ago–0.174143280.4128612Inf–0.421796196.731738e−017.052297e−01
5bon ago–chimp ago–1.983298860.4882581Inf–4.061988864.865640e−058.679249e−05
6bon ago–hum ago–5.735457800.7448880Inf–7.699758161.363240e−149.997093e−14
7bon ago–mac ago–2.371159910.3947683Inf–6.006460381.896173e−095.688518e−09
8bon ago–bon aff–0.504695070.4196264Inf–1.202724732.290829e−012.799902e−01
9bon ago–chimp aff–1.822681560.4239757Inf–4.299023581.715522e−053.431044e−05
10bon ago–hum aff–6.423269280.7143960Inf–8.991188562.445710e−199.704562e−18
11bon ago–mac aff–1.165027520.4263215Inf–2.732744226.280909e−039.421363e−03
12chimp ago–hum ago–4.514055900.6987376Inf–6.460301651.044945e−104.056844e−10
13chimp ago–mac ago–0.222307140.3099024Inf–0.717345684.731608e−015.384244e−01
14chimp ago–bon ago1.612303060.3982712Inf4.048253765.160118e−058.962311e−05
15chimp ago–chimp ago–0.196852510.2903578Inf–0.677965354.977937e−015.475730e−01
16chimp ago–hum ago–3.949011450.7019870Inf–5.625476651.849964e−084.883906e−08
17chimp ago–mac ago–0.584713570.3413709Inf–1.712839418.674209e−021.144996e−01
18chimp ago–bon aff1.281751280.4624629Inf2.771576665.578553e−038.562431e−03
19chimp ago–chimp aff–0.036235210.2699085Inf–0.134250008.932049e−019.069465e−01
20chimp ago–hum aff–4.636822930.6402440Inf–7.242274474.412218e−132.393656e−12
21chimp ago–mac aff0.621418820.3143750Inf1.976679844.807783e−026.475789e−02
22hum ago–mac ago4.291748750.7091699Inf6.051791631.432437e−094.727042e−09
23hum ago–bon ago6.126358960.6829140Inf8.970908892.940776e−199.704562e−18
24hum ago–chimp ago4.317203380.7105662Inf6.075722881.234304e−094.287583e−09
25hum ago–hum ago0.565044440.8198583Inf0.689197674.906989e−015.475730e−01
26hum ago–mac ago3.929342330.6815967Inf5.764907788.170250e−092.344507e−08
27hum ago–bon aff5.795807180.6925633Inf8.368631165.829140e−177.694464e−16
28hum ago–chimp aff4.477820680.6986217Inf6.409507101.459909e−105.352998e−10
29hum ago–hum aff–0.122767030.8810853Inf–0.139336158.891845e−019.069465e−01
30hum ago–mac aff5.135474720.6894527Inf7.448624929.431810e−145.659086e−13
31mac ago–bon ago1.834610210.4117633Inf4.455496838.369913e−061.781982e−05
32mac ago–chimp ago0.025454630.3645522Inf0.069824379.443334e−019.443334e−01
33mac ago–hum ago–3.726704310.7256292Inf–5.135824282.809100e−076.621451e−07
34mac ago–mac ago–0.362406420.2833893Inf–1.278828812.009573e−012.502488e−01
35mac ago–bon aff1.504058430.4754706Inf3.163304721.559890e−032.573818e−03
36mac ago–chimp aff0.186071930.3152714Inf0.590196065.550592e−016.005559e−01
37mac ago–hum aff–4.414515780.6599336Inf–6.689333192.241898e−119.247828e−11
38mac ago–mac aff0.843725970.3103779Inf2.718383076.560184e−039.621603e−03
39bon ago–chimp ago–1.809155580.4258084Inf–4.248755462.149614e−054.172780e−05
40bon ago–hum ago–5.561314520.6705264Inf–8.293953131.095456e−161.205002e−15
41bon ago–mac ago–2.197016630.3849704Inf–5.706975831.150011e−083.162530e−08
42bon ago–bon aff–0.330551780.3533299Inf–0.935533073.495136e−014.119268e−01
43bon ago–chimp aff–1.648538280.4027399Inf–4.093307944.252623e−057.796476e−05
44bon ago–hum aff–6.249125990.7090878Inf–8.812908981.219396e−182.682670e−17
45bon ago–mac aff–0.990884240.3686192Inf–2.688097417.186043e−031.031041e−02
46chimp ago–hum ago–3.752158940.6870573Inf–5.461201974.729216e−081.200493e−07
47chimp ago–mac ago–0.387861050.3913065Inf–0.991195013.215904e−013.859084e−01
48chimp ago–bon aff1.478603790.4980115Inf2.969015532.987555e−034.809235e−03
49chimp ago–chimp aff0.160617300.3028305Inf0.530386785.958438e−016.342853e−01
50chimp ago–hum aff–4.439970420.6631146Inf–6.695630862.147432e−119.247828e−11
51chimp ago–mac aff0.818271340.3244697Inf2.521873231.167318e−021.639212e−02
52hum ago–mac ago3.364297890.6872616Inf4.895221569.819503e−072.234784e−06
53hum ago–bon aff5.230762730.7002025Inf7.470357417.997730e−145.278501e−13
54hum ago–chimp aff3.912776240.7175861Inf5.452692314.961287e−081.212759e−07
55hum ago–hum aff–0.687811480.9013979Inf–0.763049764.454337e−015.157654e−01
56hum ago–mac aff4.570430280.6688791Inf6.832969048.317491e−123.921103e−11
57mac ago–bon aff1.866464850.4329373Inf4.311167061.623952e−053.349400e−05
58mac ago–chimp aff0.548478350.3478039Inf1.576975831.148011e−011.485661e−01
59mac ago–hum aff–4.052109360.6704754Inf–6.043636111.506791e−094.735630e−09
60mac ago–mac aff1.206132390.3120989Inf3.864583371.112790e−041.883183e−04
61bon aff–chimp aff–1.317986490.4580890Inf–2.877140724.012966e−036.306089e−03
62bon aff–hum aff–5.918574210.7345777Inf–8.057110747.811887e−167.365494e−15
63bon aff–mac aff–0.660332460.4377438Inf–1.508490661.314290e−011.668137e−01
64chimp aff–hum aff–4.600587720.6360309Inf–7.233276644.714776e−132.393656e−12
65chimp aff–mac aff0.657654040.3282083Inf2.003769994.509470e−026.200522e−02
66hum aff–mac aff5.258241750.6759408Inf7.779145937.301580e−156.023803e−14
  1. SE: standard error (of the mean); df: degrees of freedom; p-value(unc): uncorrected p-value; p-value(FDR): p-value corrected for multiple comparisons. hum: human; chimp: chimpanzee; bon: bonobo; mac: macaque; aff: affiliative context (positive for nonhuman primates; happy voices for human); ago: agonistic context (threat and distress for nonhuman primates; angry and fearful voices for human).

These results cannot be attributed to potential attentional biases toward or saliency between the calls/vocalizations of our species. Indeed, no difference in exogenous cueing was observed in an independent sample of 28 participants in a dedicated task (see Methods and Figure 1D, E).

Neuroimaging data within the sample-specific TVAs

We aimed at uncovering functional changes relative to species categorization and processing within sample-specific (N = 23) TVA, as delineated in our hypotheses. As described above, we used three distinct and independent statistical models including trial-level parametric modulators (Models 1–3). We were particularly interested in human brain activity while processing vocalizations of our closest relatives—both acoustically and phylogenetically—namely the chimpanzee but also the bonobo. The present study did not aim at uncovering whole-brain results underlying the processing of each species’ vocalizations but rather focused on human voice-sensitive areas, namely the TVA, although corrected statistics (voxelwise p < 0.05 false discovery rate [FDR]) presented in this section were computed with a whole-brain voxelwise approach for higher data reproducibility and generalizability, and not using region-of-interest (ROI) analyses. ROI analyses would most probably have artificially amplified the number of voxels in the TVA in this study. Clusters outside the bounds of the sample-specific TVA are therefore visible but in a desaturated hue to better highlight TVA activations. These clusters are more visible in figure supplements with the same contrasts as those presented in this section, but with an outline of the TVA from an independent, larger sample of participants excluding the 23 participants of this study (N = 98; Figure 2—figure supplements 13).

The three statistical models tested the impact of literature-based (Grawunder et al., 2018; Debracque et al., 2023; Gruber and Clay, 2016; Kelly et al., 2017) acoustics (1), acoustic distance between our species (2), and data-driven acoustics that represent our stimulus set the best (3). Results from all models overlap greatly, showing within-TVA enhanced activity mostly for chimpanzee calls compared to the other primate vocalizations—both when including or excluding human voice stimuli. In this section, we will focus on the latter model (3), which is the most sensitive and therefore the most meaningful for the present study. Activations coordinates can be found in Table 1, illustration in Figure 2. Data from Models 1–2 can be found in the supplementary material, see Figure 2—figure supplements 13 and Supplementary file 1C, D.

Figure 2 with 5 supplements see all
Whole-brain results with TVA outlines when contrasting the processing of chimpanzee to other species’ vocalizations with vocalization loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference as trial-level covariates of no-interest (Model 3).

Enhanced brain activity on a sagittal view with activity for [chimpanzee > human, bonobo, macaque] (dark blue to green), [chimpanzee > bonobo, macaque] (brown to red with light yellow outline), and [macaque > human, bonobo, chimpanzee] (red to yellow) vocalizations, with outlines of TVA for voice > animal sounds (A, B), voice > nature sounds (C, D), voice > music (E, F), and voice > noise (G, H). Brain activations are independent of the most discriminant low-level acoustic parameters of the stimuli set (Debracque et al., 2023). Data corrected for multiple comparisons using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. White outline: sample-specific temporal voice areas (TVA; N = 23); Dotted black outline: sample-specific TVA, per sound category; Blue outline: areas selective to chimpanzee calls. ‘a’ prefix: anterior; ‘m’ prefix: mid; ‘p’ prefix: posterior; STG: superior temporal gyrus; STS: superior temporal sulcus; L: left hemisphere; R: right hemisphere.

In addition to the abovementioned models, we also included a model-based methodology (Model 4) to uncover the brain correlates of the probability of correctly categorizing the four species, within the sample-specific TVA. This modeling includes the fitted regression coefficients computed using as covariates the acoustic parameters of Model 3, in addition to task reaction times. These neuroimaging results are presented in Figure 4.

Effects of species processing with vocalization most discriminant acoustic parameters (N = 6) as covariates of no-interest at the trial level

In this third model, we wanted to elaborate more on the discriminant factors that characterize the low-level acoustic parameters of our set of stimuli. This approach is complementary to the inclusion of both basic acoustics (Model 1) and acoustic distance in Model 2 and therefore refines these results. To do so, we used as trial-level covariates of no-interest the acoustic parameters explaining the most variance ([r > 0.7] and [r < –0.7]) in factors 1–3 of a discriminant analysis of these stimuli (Debracque et al., 2023)—see the methods section for details on this analysis. These parameters therefore include, in this specific order: vocalization loudness, intensity, change in spectrum, bandwidth contour of the second formant (F2), power of the fundamental frequency (F0), and finally the difference in intensity contour. We used these acoustic features as covariates in our data modeling. As in previous modeling of the imaging data, TVA activity was triggered by chimpanzee vocalizations ([chimpanzee > human, bonobo, macaque]) in large bilateral clusters of the aSTG within the TVA (Figure 2A–H), closely resembling activations of Model 2 (Figure 2—figure supplement 2). Chimpanzee compared to other nonhuman primate calls in this model ([chimpanzee > bonobo, macaque]) led to the largest clusters observed in the aSTG—all models considered, still within the sample-specific TVA. Indeed, we observed a large left-lateralized cluster of the aSTG extending to the mid STG (Figure 2A, C, E, G) as well as a right-lateralized cluster (Figure 2B, D, F, H). These chimpanzee-selective areas are outlined in blue in each panel of Figure 2, Table 2.

Table 2
Activations, cluster size, and coordinates for each contrast of interest of Model 3 (vocalization loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference as trial-level covariates of no-interest) in the sample-specific temporal voice areas, whole-brain voxelwise p < 0.05 FDR corrected, k > 10.
MNI coordinates
Region labelHemisphereXYZt-valueCluster size (voxels)
Chimpanzee > human, bonobo, macaque
Superior temporal gyrus ant10L–52–2–124.5472
Superior temporal gyrus ant11R540–123.1919
Chimpanzee > bonobo, macaque
Superior temporal gyrus ant13L–54–2–125.03100
Superior temporal gyrus antL–58–12–24.98
Superior temporal gyrus antL–522–124.38
Superior temporal gyrus ant14R58–2–124.9791
Superior temporal sulcus antR54–8–124.3
Chimpanzee > human
Superior temporal gyrus ant12L–50–4–122.710
Superior temporal gyrus antL–48–10–122.6
Human > chimpanzee
Superior temporal gyrus midL–54–1448.555392
Supramarginal gyrusL–60–48247.54
Superior temporal gyrus postL–54–58186.88
Middle temporal gyrus midL–66–22–86.39
Middle temporal gyrus postL–66–44–25.49
Supramarginal gyrusR56–42288.225250
Supramarginal gyrusR56–42367.52
Superior temporal gyrus midR56–827.5
Superior temporal gyrus postR50–48226.23
Superior temporal gyrus postR58–46146.22
Superior temporal gyrus midR50–1866.05
Superior temporal gyrus postR60–46185.99
Middle temporal gyrus postR70–3205.84
Middle temporal gyrus midR64–22–125.75
Middle temporal gyrus midR66–20–185.42
  1. ant: anterior; mid: central part; post: posterior.

  2. 10–14Figure S3 cluster labels.

In this last modeling of the fMRI data, no voxels reached significance either at the whole-brain level or within the TVA for the [bonobo > human, chimpanzee, macaque] or [bonobo > chimpanzee, macaque], [bonobo > chimpanzee], [bonobo > macaque] contrasts. We however found activity for the processing of macaque calls only in the left TVA, more specifically in a small cluster of the left mid STS ([macaque > human, chimpanzee, bonobo], Figure 2A, C, E, G)—and in a small portion of the planum temporale adjacent to the primary auditory cortex for the [macaque > chimpanzee, bonobo] contrast (see Figure 2—figure supplement 4). Other contrasts are reported in Figure 2—figure supplement 5.

A synthesis of the sensitivity of the human TVA to nonhuman primate calls (across statistical models)

In the previous sections, we described three different models used to analyze our fMRI data. These models—from the simplest to the more sophisticated one—highlighted enhanced activity within sample-specific bilateral anterior TVA of our participants specifically when processing chimpanzee vocalizations—but also when processing macaque calls in Model 3, in the bilateral mid STG, STS, and planum temporale. When processing chimpanzee calls, TVA activity was especially enhanced in the aSTG but also in the anterior STS. We therefore regrouped these fourteen chimpanzee-selective aSTG clusters in Figure 3—most of them overlap greatly but we still named them individually according to each contrast and analysis for exhaustivity—overlaid with sample-specific TVA (Figure 3C, D) and with the more general TVA from an independent sample of 98 participants (Figure 3A, B). Zooming closely, the area of maximal overlap between these regions (the orange surface) is located within the more general as well as within the sample-specific TVA. Interestingly, left-lateralized more medial clusters of aSTG were outside the outline of the sample-specific but not of the general TVA (Figure 3A, C), while this was not the case for right-lateralized aSTG activations. Comparing the areas recruited when processing chimpanzee to bonobo and macaque calls, this contrast—especially in Model 3, yielded distinct clusters of aSTG. This result is visible when looking at the three ‘rich blue’ outlines in every panel of Figure 3. The results synthesized here highlight the important role of acoustic parameters and emphasize the role of the most discriminant acoustic features on TVA activity relating to nonhuman primate vocalizations, especially those of chimpanzees and macaques.

Synthesis of mid-to-anterior TVA clusters of activity recruited specifically by the processing of chimpanzee and macaque vocalizations (Models 1–3).

Anterior superior temporal gyrus (aSTG) and sulcus (aSTS) clusters recruited for the processing of chimpanzee calls as opposed to human voices, bonobo, macaque calls (pink: Model 1; purple: Model 2; blue: Model 3) in the general TVA (A, B, N = 98) as well as in the sample-specific TVA (C, D, N = 23). Macaque results are only significant for Model 3 (teal: macaque vs all other species). Model 1: mean of fundamental frequency and energy (covariates of no-interest, N = 2); Model 2: acoustic distance (covariate of no-interest, N = 1); Model 3: acoustic parameters that characterize low-level acoustics of our stimuli following a discriminant analysis (covariates of no-interest, N = 6). Data are all corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05 with t-values ranging from 5 to 10. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: temporal voice areas. Prefix ‘a’: anterior; ‘m’: mid. L / R: left / right hemisphere.

These results are given even more weight by more fine-tuned comparisons of voice vs non-voice material in the voice-localizer task, namely by splitting the non-vocal blocks as a function of the auditory sounds they contain. In this more specific outline of TVA subregions, visible in Figure 2 (as defined in Figure 2—figure supplement 4), we observed that most chimpanzee- and macaque-selective STG and STS regions were still within the bounds of TVA.

Within-TVA neural correlates of the probability of correct species categorization

Determining the neural correlates of the probability of correct species categorization within the sample-specific TVA was done using a model-based approach. Specific brain correlates of behavioral data on species categorization in frontal brain regions—using different acoustic covariates modeling—can be found elsewhere (Ceravolo et al., 2023).

The present data revealed massive TVA activations when including all species together (including human), namely in the posterior, mid, and anterior STG, STS, and middle temporal gyrus (MTG; Figure 4A, B). Looking separately at these neural correlates for each nonhuman primate species, results show smaller, more specific areas for chimpanzee vocalizations, namely bilateral mSTG and mSTS in addition to right mMTG (Figure 4C, D). For bonobo vocalizations, large correlates were observed bilaterally, covering most of the TVA. More specifically, the strongest peaks were located in the bilateral mSTG, mSTS, and mid and posterior MTG (Figure 4E, F). Finally, for macaque vocalizations, correlates were solely observed in a small cluster of the right mSTS (Figure 4G, H). According to these data, correctly categorizing our species stimuli was associated with strong and vast TVA regions for bonobo calls, corresponding to the species with the lowest accuracy of all—and the only species with below-chance average values.

Model-based correlates of the probability of correct species categorization, within sample-specific TVA (Model 4).

Correlates of the probability of correct species categorization computed using model-based analysis technique for all species, as illustrated on sagittal renders for all species, including human (A, B), then specifically for chimpanzee calls (C, D), bonobo calls (E, F), and macaque calls (G, H). These correlates were constrained to the bounds of the sample-specific TVA (N = 23, black outline) using an inclusive masking procedure with correction for multiple comparison using voxelwise false discovery rate (FDR) at a threshold of p < 0.05. The colorbars represent t-value statistics. TVA: temporal voice areas. Prefix ‘a’: anterior; ‘m’: mid; ‘p’: posterior. STG: superior temporal gyrus; STS: superior temporal sulcus; MTG: middle temporal gyrus; PT: planum temporale. L/R: left/right hemisphere.

Acoustic differences between chimpanzee and bonobo calls

The differences observed between chimpanzee and bonobo stimuli, both at the behavior and brain levels, may originate in their acoustic specificities—while phylogeny relative to humans is almost identical. While we will discuss this aspect in the Discussion, we explored it using a separate acoustic analysis dedicated to highlighting such acoustic differences. This analysis is different from the one in Model 3—although the same methods were used, because here we only use as input for GeMAPS (Eyben et al., 2015; Eyben et al., 2010) the calls from chimpanzees and bonobos, therefore constraining acoustic characterization to these two species. The results highlight the acoustic parameters that explain the most variance for these specific species (see Methods and Supplementary file 1E). The reduced parameters explaining at least 80% of the acoustic variance between chimpanzee and bonobo calls relate to (in that order): (1) pitch variation within the call (‘F0semitoneFrom27.5Hz_sma3nz_pctlrange0.2’), (2) within-call variability in mid-frequency range energy, sensitive to changes in voice quality and resonance (‘slopeV500.1500_sma3nz_stddevNorm’), (3) vocal tract opening (‘F1frequency_sma3nz_stddevNorm’), (4) irregularity in the temporal structure of unvoiced intervals (‘StddevUnvoicedSegmentLength’), (5) ratio of periodic to aperiodic energy (voice regularity, breathiness; ‘HNRdBACF_sma3nz_stddevNorm’), (6) spectral slope between 0 and 500 Hz in voiced frames (‘creaky’ or ‘pressed’ phonation; ‘slopeV0.500_sma3nz_amean’), (7) instability in glottal closure quality (‘F1bandwidth_sma3nz_stddevNorm’), (8) steepness of pitch rise between consecutive voiced frames (‘F0semitoneFrom27.5Hz_sma3nz_stddevRisingSlope’), (9) changes in place of articulation or lip posture (‘F2amplitudeLogRelF0_sma3nz_stddevNorm’), (10) within-call changes in spectral shape of very low frequency (‘slopeV0.500_sma3nz_stddevNorm’), (11) lip rounding and vocal tract length (‘F3amplitudeLogRelF0_sma3nz_stddevNorm’), (12) Jitter variability (micro-irregularities in vocal fold vibration; ‘jitterLocal_sma3nz_stddevNorm’), and (13) spectral tilt measure (within-call changes in voice quality; ‘logRelF0.H1.A3_sma3nz_stddevNorm’).

Discussion

The present study provides evidence of the sensitivity of the human TVA to cross-species vocalizations, especially to chimpanzee calls but also to macaque vocalizations, as illustrated by selective enhanced activity in the bilateral mid and anterior STG and STS—within sample-specific TVA. We do not interpret these results as ‘absolute’ selectivity, but rather we emphasize the relative preference of these brain regions to chimpanzee and to a lesser extent macaque calls. These results were obtained through statistical modeling of the MRI data that included either simple acoustics or the use of Mahalanobis acoustic distance between species and the most discriminant acoustic features specific to our stimuli as covariates. These two latter analyses converged and yielded greatly overlapping results, especially in the anterior TVA. Therefore, our results suggest that vocalizations from another ape species recruit, or trigger activity in, subregions of human temporal cortex that process species-specific voices in humans—namely the bilateral, sample-specific TVA. This evidence speaks in favor of cross-species primate vocalization processing in the anterior and mid TVA of humans—for chimpanzee and macaque calls, respectively. Our results also suggest differential engagement of the TVA according to behavioral performance for species categorization. While our acoustic data confirmed the hypothesized hierarchy of acoustic distance as a function of phylogenetic distance between our species, we still observed mid STG and STS activity for macaque vs bonobo calls and a small cluster in the left mid STS selective to macaque calls in Model 3—an unexpected result since macaques are the most distant species from humans both phylogenetically and acoustically in our study. Therefore, while we initially hypothesized that primate calls would exclusively enhance activity in human TVA as a function of a combination of phylogenetic and acoustic proximity, our data also point toward the greater importance of the most discriminant acoustic features rather than acoustic distance alone. We discuss these aspects below in more detail and interpret their general meaning and subsequent scientific implications, in addition to highlighting the limitations of our study.

Often specifically associated with the processing of conspecific vocalizations (e.g., in humans Belin et al., 2000; Fecteau et al., 2004; Bodin et al., 2021), macaques (Petkov et al., 2008; Bodin et al., 2021; Perrodin et al., 2015), and dogs (Andics et al., 2014), the present study challenges the common view of the TVA as ‘species-specific’ and illustrates that human voices, chimpanzee, and macaque calls can enhance activity in the TVA. We think that the distinct locations of these TVA subregions recruited for processing the vocalizations of these primate species matters. In fact, there might be a possible association between anterior TVA—selective to processing chimpanzee calls—and the higher recognition performance of chimpanzee calls compared to those of bonobos or macaques in human participants (Ceravolo et al., 2023). Anterior TVA activity selective to the processing of chimpanzee calls occurred when these were compared to both human and nonhuman primate species, solely to other nonhuman primate vocalizations, or directly to the human voice. However, homologous results were not observed for bonobo, and they were more scarce—especially between models—for macaque vocalizations: we found macaque-selective activity in a small area of the planum temporale and in a small cluster of the left mid STS, congruent, for instance, with locations observed in the general processing of animal sounds, especially in the planum temporale (Engel et al., 2009; Lewis et al., 2005). On the other hand, within-TVA anterior STG activity was also observed when chimpanzee vocalizations were directly compared to the human voice. We think this result highlights the cross-species selectivity of this anterior subregion of the TVA for processing species phylogenetically close to humans and especially with human-like acoustics, namely the calls of chimpanzees in the case of our study. Because of their vocal proximity, the perception of human voices and chimpanzee calls in socio-affective contexts could involve a common ‘social’ core of the brain, which increases activity in brain regions such as the anterior TVA, as reported previously in studies pertaining to social contextual information processing in the anterior STG (Simmons et al., 2010; Zahn et al., 2007). Differences at the level of processing complexity between the two types of vocalizations could also explain this observation, while we demonstrated that saliency or attention-related effects do not exist between our species stimuli. Indeed, previous studies have shown the role of the anterior STG and the anterior STS in the conceptual representation of social context by the human voice (Simmons et al., 2010; Zahn et al., 2007; Bestelmeyer et al., 2011; Mellem et al., 2016). Therefore, our data might suggest that the anterior part of the superior temporal cortex could be recruited to process the social context of human and chimpanzee vocal stimuli. However, this processing would be more automated for the perception of the human voice than for chimpanzee calls because of our high exposure and expertise as humans to these vocal signals, but these hypotheses should be addressed scientifically in studies dedicated to this topic. Our results are also complementary to and coherent with a ‘voice patch’ system in the brain of primates, as put forward by Belin et al., 2018, and according to which distinct ‘patches’ or subregions of the temporal lobe—especially its anterior portion—would be interconnected and would allow for the processing of voice information. Such a system would be present in many primate species such as humans, macaques, and marmosets, with most recent evidence suggesting a population of neurons in the anterior STG of the macaque brain selective to human voice (Giamundo et al., 2024), as also anticipated in another study on macaques and also in the anterior STG (Perrodin et al., 2014). These fascinating and converging results mirror our present data—with chimpanzee calls triggering responses in the anterior STG/TVA of our human participants—and strongly emphasize the need for pursuing a comparative approach in order to clarify the cross-species neural bases underlying the processing of human and nonhuman primate vocal signals. As we mentioned previously, these interpretations are, considering our results, free of any potential attentional bias toward one species over the others, since no effect was observed on that matter in a control, behavioral study involving an independent sample of 28 participants in a species-specific exogenous cueing attentional paradigm—Methods and Figure 2—figure supplement 1.

Importantly, our data also emphasize the influence of acoustic features and especially acoustic proximity between human and chimpanzee vocalizations: we show that activity in the anterior STG and more generally in the anterior TVA partly depends on phylogenetic and, importantly, on key acoustic features—including acoustic proximity. Consistent with previous studies (Belin et al., 2008a; Fritz et al., 2018), we did not expect TVA activity for macaque calls processing because they are both phylogenetically and acoustically more distant from humans than the other species in this study—although, as mentioned above, we found a very small cluster in the left primary auditory cortex and mid STS for Model 3. It is interesting to note that in Model 2, with acoustic distance as trial-level covariate, we observed TVA activity only for chimpanzee but not macaque calls, giving further weight to the importance of acoustic distance in this context. On this matter, the mid STS location of the macaque-related cluster aligns with a cortical region preferentially responsive to non-rhythmic, ‘noise-like’ sounds—a category that acoustically includes macaque vocalizations—rather than to harmonic, voice-like signals (Norman-Haignere et al., 2019). This is further consistent with the posterior-to-anterior functional gradient of the superior temporal lobe, whereby mid and posterior regions process lower-level acoustic features and more anterior STG/STS regions support categorical sound recognition (Leaver and Rauschecker, 2010). Also, if only phylogenetic proximity mattered, bonobo calls should also elicit activity in the TVA because they are as phylogenetically close to humans as chimpanzees. But this viewpoint is rather reductive, and our results show that this is not entirely correct, and that activity in the TVA crucially depends on the acoustic properties of the perceived vocalizations since we cannot infer phylogeny from vocalizations. This interpretation is strongly supported by the inclusion of acoustic Mahalanobis distance for each species compared to the human voice as a trial-level covariate of no-interest. Using such modeling, differential neuroimaging results between chimpanzee and bonobo vocalizations were explained by both acoustic and phylogenetic proximity in the TVA. These results are consistent with the recent proposal—and recent findings (Debracque et al., 2023)—that there are substantial differences between chimpanzee and bonobo vocalizations. These encompass fundamental frequency range and mean due to larynx length (Grawunder et al., 2018; Titze et al., 2016)—despite the evolutionary relatedness to chimpanzees (Grawunder et al., 2018). Therefore, the interaction between phylogeny and acoustic distance or proximity would explain the anterior TVA expansion for processing specifically chimpanzee but not bonobo vocalizations. This argument, however, falls short to explain the activations of the TVA by macaque calls in Model 3, that may have been treated as less complex or structured sounds in the mid TVA (Norman-Haignere et al., 2019).

Overall, it seems reasonable to hypothesize that TVA activity is not per se human-specific (Belin et al., 2000; Bestelmeyer et al., 2011), but that TVA are instead sensitive to vocalizations from other primate species, provided that these vocalizations have sufficient acoustic proximity to human vocal signals—which would in itself be related to anatomical and/or behavioral changes throughout phylogenetic evolution. This integrative view is again consistent with the concept of a ‘voice patch’ system in the primate brain (Belin et al., 2018). We therefore propose that the mid and anterior TVA, unlike the rest of the TVA, would be heterospecific—sensitive to vocalization acoustics triggered by evolution. This proposition also implies a validation of our mechanistic hypothesis according to which the mean fundamental frequency of chimpanzee but not bonobo calls—the former being much closer to the mean fundamental frequency of the human voice (Grawunder et al., 2018), would allow for a better identification and recognition of chimpanzee calls by humans. This advantage would rely on neurons of the human auditory cortex—both the primary and more secondary regions—being specialized in the processing of low to mid fundamental frequencies such as those of the human voice and chimpanzee calls. In our third analysis model, we looked further into this aspect and included several acoustic properties of our stimuli as a function of the four species in our stimuli. A discriminant analysis (Debracque et al., 2023) allowed us to select specific acoustic features that best discriminate between our species stimuli. Namely, we took the six parameters explaining the most the differences between our stimuli, including vocalization loudness, intensity—similar to our ‘energy’ covariate of Model 1, in addition to change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference. Using these more sophisticated acoustic features as covariates of no-interest, we still obtained brain imaging results very similar to those of Model 1 and even closer to Model 2—with acoustic distance as covariate, yet with some subtle differences in anterior STG cluster size and location. The peaks were indeed located more ventral and were larger—as compared to results of Models 1 and 2, especially for the processing of chimpanzee- and macaque-selective vocalizations compared to all primate species and to nonhuman primates alone. These results suggest that the inclusion of spectrum change to intensity- and frequency-related acoustical parameters of the vocal signals slightly shifted and enlarged activation locations in the anterior STG. This result is again congruent with the proposed existence of ‘voice patches’ in the temporal lobe of primate species (Belin et al., 2018), with the interconnectivity of these patches highly depending on very fine-grained acoustic aspects of primate vocal signals. This step motivated the inclusion of these parameters as covariates of no-interest in neuroimaging Model 3, to retain brain activations marginally independent of such acoustics. The congruence between these data should be explored in more detail in the future by the combination of computational bioacoustics and functional neuroimaging, due to the high relevance and sensitivity of combining these techniques to investigate primate social communication (Cauzinille et al., 2024).

A final but maybe more secondary interpretation arising from our results regarding bonobo calls also supports the evolutionary divergence of this peculiar species. According to the self-domestication hypothesis, bonobos would have evolved differently than chimpanzees due to selection against aggression (Hare et al., 2012). Interestingly, differentiation in the evolutionary path of bonobos has influenced both their behavior (Gruber and Clay, 2016) and morphology, leading to differences at the level of call production (Grawunder et al., 2018; Kelly et al., 2017). Considering these documented acoustic differences and putting them in perspective with our neuroimaging data, the calls of our last common ancestor with the other Pan species 8 million years ago (Perelman et al., 2011), may have been closer to those uttered by modern chimpanzees than to those of bonobos. Our data indeed show that modern human brains remain more sensitive to the acoustic characteristics of the calls of the former compared to the latter, arguing for more conserved calls between modern chimpanzees and humans. This aspect is also in line with significant differences between, for instance, the fundamental frequency of human baby cries or babbling (~250–600 Hz) compared to that of bonobos (~1000–3500 Hz) (Dewi et al., 2019; Jagtap et al., 2016), while they correspond more closely to the fundamental frequency of chimpanzee calls (~500–1000 Hz) (Grawunder et al., 2018). In our study, bonobo calls definitely are so much different than those uttered by the species of our other stimuli, that they presumably fall outside of the phylogeny and acoustic proximity factors that we outlined so far. This would also put into perspective the selectivity of mid TVAs for macaque calls. The distinct analysis we computed, dedicated to acoustic differences between chimpanzee and bonobo calls, highlighted specificities in pitch dynamics, phonation quality—periodicity, breathiness, glottal irregularity—and vocal tract resonance. This analysis therefore revealed the more stable, lower-pitched, periodically structured nature of the calls of chimpanzees relative to the higher, breathier, and more variable calls of bonobos. These data align well with those obtained through Mahalanobis acoustic distance data discussed above, again converging toward voice-like acoustics for chimpanzee calls.

In a sense, we therefore validate our first hypothesis regarding the existence of acoustic distance between each primate species used in our study. We also partially validate our second hypothesis, albeit not completely. In fact, macaque compared to bonobo or other primate calls in Model 3 revealed mid TVA activations, and we think that these activations may depend specifically on the importance of the most discriminant acoustic features. Several TVA subregions or ‘patches’ underlying cross-species primate vocalization processing might therefore exist, and our data highlight at least one of them in the mid and anterior portion of the TVA.

Behavioral data should also be discussed in the context of the species categorization task. We know from our attentional control task that species categorization was not biased by greater saliency from a specific species. We can therefore consider the proportion of correct categorization as attention-bias free in that regard. A ceiling effect was expected and observed for human voice, while only chimpanzee and macaque vocal signals were categorized above chance level for nonhuman primate species. Interestingly, no difference was observed in accuracy between these two, while bonobo calls were categorized below chance. These results are at odds with the acoustic distance argument per se, since we show that macaque coos are the furthest from human voice and yet these were better recognized than acoustically closer bonobo calls. The neural correlates of species categorization performance also point toward the calls of bonobos being completely different than the other species, recruiting most TVA territory compared to other species. Combined, and aligning with previous data (Debracque et al., 2023; Ceravolo et al., 2023; Debracque et al., 2025), these puzzling behavioral and neural results illustrate the very different nature of bonobo calls, and the fact that acoustically, they may have even represented an ‘object’ that was not part of the primate category for our participants. Considering the vast TVA activations compared to the other species, bonobos calls were also most probably very engaging, congruent with massive resources being allocated due to hesitation and decisional difficulty (Ceravolo et al., 2023).

Considering behavioral data and acoustics, we should also mention the possibility that familiarity to some species may have influenced our results. In fact, differential familiarity with the stimulus species through media or zoo exposure cannot be entirely ruled out. However, this interpretation is undermined by the observation that bonobo calls failed to elicit any anterior TVA response. This interpretation also suffers from the fact that bonobos are phylogenetically equidistant from humans as chimpanzees, and are equally unfamiliar to Western participants, although maybe less culturally present than chimpanzees. Consistent with this nuanced viewpoint, ERP evidence suggests that early, automatic neural responses to primate vocalizations are modulated by phylogenetic proximity rather than familiarity—with the latter influencing only later attentional processing stages (Scheumann et al., 2014; Scheumann et al., 2017).

A fundamental challenge, shared by vocal science, cognitive and affective neuroscience fields, relates to dissociating biologically relevant information from acoustic properties: species-specific vocal signals are in fact entirely instantiated in acoustics—there is no ‘non-acoustic’ biological residual in the strict sense (Debracque et al., 2023; Cauzinille et al., 2024; Leaver and Rauschecker, 2010). Our Models 1–3 therefore demonstrate independence from specific, theory-motivated parameterizations of the acoustic signal rather than from acoustic information in its entirety. Therefore, residual anterior TVA response to chimpanzee calls most plausibly reflects sensitivity to higher-order acoustic structure not captured—at least not entirely—by these lower-dimensional features. This higher-order acoustic structure may include patterns of spectro-temporal properties of species-typical calls, their sequential organization, or the results could also depend on an evolutionary ‘tuning’ of the human anterior TVA to an acoustic processing niche shared with the great ape vocal system (Belin et al., 2018). We acknowledge this as a fundamental limitation of the acoustic-control approach in comparative vocal neuroimaging, and in our study, and suggest that encoding models using richer spectro-temporal representations—such as, for instance, more complex sequences of calls rather than single bursts—may offer more exhaustive acoustic control in future work.

Therefore, even though we tried to control for critical acoustic features, species categories, and their related evolutionary distance, several limitations should be mentioned. These limitations are both theoretical and methodological. First, we cannot rule out the fact that including more primate species in our set of stimuli would have influenced the results. In fact, even though our species categories were specifically chosen for this task, the inclusion of vocalizations from other great apes—such as gorillas or orangutans—would have broadened the scope of our results. Related to this aspect, we can also mention that tackling primate phylogeny, which spans over millions of years, with only four species restricts the possible inference based on our results. Also, not including either nonvocal stimuli or scrambled nonhuman primate calls limits the interpretability of the results, since we cannot rule out that the auditory species ‘object’ per se was triggering the results. Active species categorization compared to passive listening may also introduce a decision-making bias in the neural effects observed. In fact, we observed improved sensitivity of our data by the use of more sophisticated acoustic modeling, namely the inclusion of both between-species acoustic distance and of the most discriminant acoustic features in the functional imaging data. However, we did not include as stimuli—or in a control task—the synthesized acoustic parameters of interest, for instance, by using species-specific F0 contour or its spectral content in other neutral, comparable auditory stimuli. We cannot therefore completely rule out that such tasks would not trigger brain activations that overlap with our results—although such data would not be mutually exclusive with our data and interpretation. Future work should therefore address with the greatest level of detail the specific question of acoustics in primate vocalization processing, in addition to adding more—as well as synthesized—stimuli from other great ape species, and if possible using both bursts as well as more complex sequences of calls. The origin of these acoustic differences should also be investigated, since we can assume that these differences originated at least partially from evolutionary processes as well as survival and adaptation mechanisms. Linked to this aspect, phylogeny and acoustics were at least partly confounded in our factorial design, and using dedicated synthetic vs natural vocal material as stimuli would disentangle this issue. Setting the acoustic covariates to zero in the fMRI models also did not inform on the impact of having acoustics of interest, and this aspect should be studied in the future. Finally, individual differences in the processing and preference of one species over another or over all the others cannot be ruled out, even though we provide evidence that attentional effects toward the vocalizations of a specific species did likely not exist in our data. Therefore, individual differences should be assessed in more detail in the future, with the inclusion of participant-level covariates such as questionnaire scores assessing the familiarity with primate vocalizations or the hedonic value of these vocalizations for each individual. Among the more general limitations of nonhuman primate neuroscience lies the fact that more inclusive and large-scale collaborations would be needed. Such collaborations and frameworks would lead to a better study and understanding of primate neuroscience, and previous initiatives have recently been put forward in this direction (Hartig et al., 2023; Milham et al., 2022).

Taken together, our data suggest that phylogeny-driven specific acoustic features appear to be necessary to trigger cross-species activity in the human TVAs—especially in subregions of the TVA in which increased activity underlies voice signals compared to animal and nature sounds. We provide evidence for specific anterior and mid TVA subregions that underlie the processing of the calls of one of our closest relatives, namely chimpanzees but also of macaques, respectively. In line with recently reported literature, we contend that the human TVA are also involved in the processing of heterospecific primate vocalizations, provided they exhibit sufficient phylogenetic and especially spectro-temporal acoustic proximity to the human voice; as such, we predict that other similarities will be uncovered in the processing of human and nonhuman primate communicative signals. Finally, our results support a critical evolutionary continuity between the structure of human and chimpanzee vocalizations, possibly reflecting one of their common ancestors, as opposed to bonobo vocalizations that underwent more recent and critical changes within the last 1–2 million years. In contrast, the chimpanzee vocal system may be closer to that of the common ancestors of humans and chimpanzees, as shown by the conserved activation in the human modern brain.

Materials and methods

Species categorization task

Participants

Twenty-five right-handed, healthy, either native or highly proficient French-speaking participants took part in the study. One participant was excluded because he had no correct response at all and may have fallen asleep, while another participant was excluded due to incomplete scanning and technical issues at the MRI scanner, leaving us with 23 participants (10 female, 13 male, mean age 24.65 years, SD 3.66). With this sample size and our study design, we achieved a power of 75.12% for a between-means comparison with effect size dz = 0.5 and alpha = 0.05 as calculated in G*Power version 3.1.9.7 (Kang, 2021). All participants were naive to the experimental design and study, had normal or corrected-to-normal vision, normal hearing, and no history of psychiatric or neurologic incidents. Participants gave written informed consent for their participation in accordance with ethical and data security guidelines of the University of Geneva. The study was approved by the Ethics Cantonal Commission for Research of the Canton of Geneva, Switzerland (CCER) and was conducted according to the Declaration of Helsinki.

Stimuli

Request a detailed protocol

Seventy-two vocalizations of four primate species (human, chimpanzee, bonobo, and rhesus macaque) were used in this study (see Figure 1A). The eighteen selected chimpanzee, bonobo, and rhesus macaque vocalizations contained single calls or call sequences produced by six to eight different individuals in agonistic (threat, distress) or affiliative (‘positive’) social contexts. These were randomly selected—in a between-participants fashion—among our full database of primate stimuli, containing specifically: 15 chimpanzee individuals (recorded in the wild in the Budongo forest, Uganda), 10 bonobo individuals (recorded in the wild in the Salonga National Park, Democratic Republic of Congo, DRC), and 16 macaque individuals (recorded in the wild from semi-free monkeys on Cayo Santiago, Puerto Rico). We then selected eighteen human voices obtained from a nonverbal validated stimuli set of Belin and collaborators (Belin et al., 2008b), which were expressed by two male and two female adults expressing positive or negative social interactions. All vocal stimuli were standardized to 750 ms using PRAAT (https://praat.org/) but were not normalized in any way in order to preserve the naturalness of the sounds (Ferdenzi et al., 2013) and to allow for low-level acoustic parameters to be used in neuroimaging data modeling.

Experimental procedure and paradigm

Request a detailed protocol

Laying comfortably in a 3T scanner, participants listened to a total of 72 stimuli pseudo-randomized and played binaurally using MRI compatible earphones at 70 dB SPL (Model ‘S14’, Sensimetrics Corporation, Gloucester, MA, USA). At the beginning of the experiment, participants were instructed to identify the species that expressed the vocalizations using a keyboard. Three vocalizations of each species were first presented outside the scanner to the participants, revealing to them which species the stimuli were from—as a form of familiarization to the auditory material and the species. These stimuli were then discarded from the task. For instance, the instructions could be ‘Human—press 1, Chimpanzee—press 2, Bonobo—press 3, or Macaque—press 4’. The pressed keys were pseudo-randomly assigned across participants (response box: fORP, Cortech Solutions, Inc, Wilmington, NC, USA). In a 3- to 5-s interval (jittering of 400 ms) after each stimulus, participants were asked to categorize the species. If the participant did not respond during this interval, the next stimulus followed automatically. See Figure 1A for a detailed illustration of the paradigm.

TVAs localizer task

Participants

Two independent samples of participants performed the task while undergoing fMRI scanning: the sample of this study (10 female, 13 male, mean age 24.65 years, SD 3.66), leading to the delineation of sample-specific TVA; an independent sample of 98 right-handed, healthy, either native or highly proficient French-speaking participants (52 female, 46 male, mean age 24.66 years, SD 4.97) leading to the delineation of more general and representative TVA. All participants were naive to the experimental design and study, had normal or corrected-to-normal vision, normal hearing, and no history of psychiatric or neurologic incidents. Participants gave written informed consent for their participation in accordance with ethical and data security guidelines of the University of Geneva. The study was approved by the Ethics Cantonal Commission for Research of the Canton of Geneva, Switzerland (CCER) and was conducted according to the current regulations in Switzerland.

Stimuli and paradigm

Request a detailed protocol

Auditory stimuli consisted of sounds from a variety of sources (Belin et al., 2000). Vocal stimuli were obtained from 47 speakers: 7 babies, 12 adults, 23 children, and 5 older adults. Stimuli included 20 blocks of vocal sounds and 20 blocks of non-vocal sounds. Vocal stimuli within a block could be either speech 33%: words, non-words, foreign language, or non-speech 67%: laughs, sighs, various onomatopoeia. Non-vocal stimuli consisted of natural sounds 14%: wind, streams, animals, 29%: cries, gallops, the human environment, 37%: cars, telephones, airplanes, or musical instruments, 20%: bells, harp, instrumental orchestra. The paradigm, design, and stimuli were obtained through the Voice Neurocognition Laboratory website (http://vnl.psy.gla.ac.uk/resources.php). Stimuli were presented through earphones (Model ‘S14’, Sensimetrics Corporation, Gloucester, MA, USA) at an intensity that was kept constant throughout the experiment 70 dB sound-pressure level. Participants were instructed to actively listen to the sounds. The silent inter-block interval was 8 s long.

Exogenous attention task using species as cues

This task was performed to control for potential attentional biases or specific salience of one species compared to the others. This task was therefore a control task, and the results are reported in Figure 1. The task was designed according to work on attention orienting following vocal material presentation (Ceravolo et al., 2016). More specifically, a cue is first presented and is quickly followed by the presentation of a neutral target to detect. The detection of this target can reliably be delayed or accelerated depending on the nature of the cue. In this control study, the cue corresponded to the stimuli of the main species categorization task followed by a sine wave tone (a ‘bip’) that had to be detected as fast as possible (the target). Any specific attentional bias of a species would therefore trigger reaction time differences for target detection, while no differences would invalidate any attentional effect or increased salience linked to a specific species. We did not observe any difference between cue species in this task (χ2(2) = 3.33, p = 0.34, Figure 1).

Participants

Twenty-eight participants (independent of the samples presented so far) took part in this behavioral study (15 female, 13 male, mean age 22.63 years, SD 5.00). With this sample size and our study design, we achieved a power of 82.48% for a between-means comparison with effect size dz = 0.5 and alpha = 0.05 as calculated in G*Power version 3.1.9.7 (Kang, 2021). All participants were naive to the experimental design and study, had normal or corrected-to-normal vision, normal hearing, and no history of psychiatric or neurologic incidents. Participants gave written informed consent for their participation in accordance with ethical and data security guidelines of the University of Geneva. The study was approved by the Ethics Committee of the University of Geneva (Switzerland), Department of Psychology and Educational Sciences, and was conducted according to the Declaration of Helsinki.

Stimuli and paradigm

Request a detailed protocol

The cues of the task were the species stimuli (human voice, chimpanzee, bonobo, and macaque calls) used in the species categorization task above. The target ‘bip’ was a sine wave tone created using Matlab (Matlab 2020a, The Mathworks, Inc, Natick, MA, USA) with a wave frequency of 600 Hz, a fade-in and fade-out of 10ms each, and a total duration of 100ms. The auditory material was presented through headphones (Sennheiser HD-25 II, Sennheiser electronic SE & Co KG, Germany) at a constant sound-pressure level of 70 dB. The procedure unfolded as follows, on a computer with a light grey screen background: for each trial (N = 172), a cue (human voice, chimpanzee, bonobo, or macaque call) was presented for 750 ms while a black fixation cross was presented at the center of the screen. Following a jittered blank screen of duration 100–250 ms (in steps of 50 ms), the target bip was presented for 100 ms. Right at the end of the presented bip, the fixation cross turned white, indicating that the response screen had started and that a response was expected as fast and accurately as possible as instructed, by using the ‘space’ key of the keyboard. A varying inter-trial interval of 1–2.5 s was used (in steps of 500 ms). Among the total of 172 trials, there were 24 trials per species (12 in agonistic and 12 in affiliative social contexts) for a total of 96 stimuli (4 species * 24 trials = 96) that were each presented twice (N = 96*2 = 172).

Behavioral data analysis

Species categorization task

Request a detailed protocol
Accuracy
Request a detailed protocol

Behavioral data were used to exclude participants who had below chance level categorization of human voices. Therefore, data from 23 participants mentioned in the Species Categorization Task—Participants section above were analyzed using R Studio software (R Studio Team Inc, Team, R, 2015, Boston, MA, url: http://www.rstudio.com/) and the lme4 (Bates, 2014) and emmeans (Lenth, 2019) packages. These data can be found in a published article focused on decisional aspects in the frontal cortex—ROI analysis and computational modeling of the probability of correct species categorization—using the same species stimuli as in this study, but with different modeling of acoustics (Ceravolo et al., 2023).

For the present study, behavioral data were analyzed in a similar fashion as previously published by Ceravolo et al., 2023 but with acoustic covariates of neuroimaging Model 3, as described below (vocalization loudness, intensity, change in spectrum, bandwidth contour of the second formant, power of the fundamental frequency, and the difference in intensity contour). Namely, the dependent variable was a binarized accuracy for each trial (response), determining whether the participants categorized the species correctly or incorrectly (1 or 0, respectively). A logistic regression was used to predict this binary response variable with interacting Species and vocalization Context, in addition to the main effect of each of the 6 most discriminant acoustic parameters as fixed effects. Random effects included the identity of the participants (ParticipantID), response reaction times (RT)—as these were of no interest but worth controlling for, participant gender (Gender), and age (Age), and stimulus identity (StimID). This model explained 62.10% of the variance in the data (R2c).

The formula, subsequently referred to as neuroimaging Model 4, was the following:

species categorization task model<glmer(responseSpeciesContext+(Loudness+SoundIntensity+ChangeInSpectrum+F2BandwContour+F0PowerH1_H2+IntensityContDiff)+(1|ParticipantID)+(1|RT)+(1|Gender)+(1|Age)+(1|StimID))

Exogenous attention task

Reaction times

Request a detailed protocol

In this study, the dependent variable of interest was the reaction times to detect the target ‘bip’. In order to remove extreme values, we discarded for each participant the values below the 5th percentile and above the 95th percentile. On average, the number of trials per participant therefore went down from 172 to 165 (~4.1% of trials removed). Data were then analyzed using R Studio software (R Studio Team Inc, Team, R, 2015, Boston, MA, url: http://www.rstudio.com/) using linear mixed effects modeling of the lme4 package (Bates, 2014). The formula was the following:

attentional control task model<lmer(RTSpeciesContext+(1+StimID|ParticipantID)+(1|Gender)+(1|Age))

in which RT is the reaction times (dependent variable), Species is the four species of the stimuli (human, chimpanzee, bonobo, macaque) interacting with the Context of production (affiliative, agonistic) (fixed effects). The random effects were: the random slope of the identity of each stimulus (StimID) as a function of participants (ParticipantID), participant gender (Gender), and age (Age). This model explained 57.54% of the variance in the data (R2c).

There was no significant effect of species (χ2(3) = 3.33, p = 0.34), a significant effect of context (χ2(2) = 11.39, p < 0.01), and no interaction between species and context (χ2(6) = 3.06, p = 0.80). The effect of Context was explained by slower reaction times for affiliative than agonistic vocalizations, independent of species (χ2(1) = 6.08, p < 0.05). Descriptive statistics per species for the reaction times were the following: Human, mean = 207.28, SD = 118.43; Chimpanzee, mean = 211.31, SD = 127.07; Bonobo, mean = 209.45, SD = 151.85; Macaque, mean = 214.05, SD = 117.84. Descriptive statistics per context for the reaction times were: Affiliative, mean = 215.03, SD = 115.41; Agonistic, mean = 208.16, SD = 135.39. See illustration in Figure 2—figure supplement 1.

Acoustic analysis of the vocalizations

Mahalanobis acoustic distance (neuroimaging Model 2)

Request a detailed protocol

To quantify the impact of acoustic similarities in human recognition of affective vocalizations of other primates, we extracted 88 acoustic parameters from all vocalizations using the extended Geneva Acoustic parameters set defined as the optimal acoustic indicators related to voice analysis (GeMAPS Eyben et al., 2015). This open-source set of acoustical parameters was selected based on (1) their potential to index affective physiological changes in voice production, (2) their proven value in former studies as well as their automatic extractability, and (3) their theoretical significance. GeMAPS relies on an automatic extraction system, which therefore automatically extracts an acoustic parameter set from an audio file—in an unsupervised, minimalistic manner. Then, to assess the acoustic distance between vocalizations of all species, we ran a GDA model. More precisely, we used the 88 acoustical parameters in a GDA in order to discriminate our stimuli based on the different species (human, chimpanzee, bonobo, and rhesus macaque). Among these 88 acoustical parameters, we excluded those that were strongly correlating—that is., with correlation scores r > 0.70—to avoid redundancy and minimize multicollinearity. Following this selection process and the GDA, we eventually retained 16 acoustic parameters (Supplementary file 1A). Namely, the parameters were, in that order of importance: (1) Arithmetic mean of the contour of the melodic-frequency cepstral coefficient 4; (2) Arithmetic mean of the contour of the frequency with most flux around it (change in spectrum); (3) Normalized standard deviation of the contour of the frequency with most flux around it (change in spectrum); (4) Conversion of power to decibel of the root mean square of the sound energy; (5) Arithmetic mean of the slope of timeframe 0–500; (6) Arithmetic mean of the contour of the melodic-frequency cepstral coefficient 1; (7) Arithmetic mean of the second formant amplitude power log relative to fundamental frequency; (8) Normalized standard deviation of the contour of the second formant amplitude power log relative to fundamental frequency; (9) Arithmetic mean of power log relative to fundamental frequency of the harmonic one and two; (10) Normalized standard deviation of the contour of the relative amplitude difference in local dB; (11) Arithmetic mean of the contour of the sum of the auditory spectrum; (12) Rate of peaks per second of the sum of the auditory spectrum; (13) Normalized standard deviation of the contour of the first formant amplitude power log relative to fundamental frequency; (14) Arithmetic mean of the contour of the bandwidth of the second formant; (15) Arithmetic mean of the contour of the relative amplitude difference in local dB; (16) Normalized standard deviation of the contour of the melodic-frequency cepstral coefficient 2.

We subsequently computed multi-dimensional Mahalanobis distances to classify the 72 stimuli on these selected acoustical features. A Mahalanobis distance is a generalized pattern analysis comparing the distance of each vocalization from the centroids of the different species vocalizations. This analysis allowed us to obtain an acoustical distance matrix used to test how the acoustical distances were differentially related to the different species (see Figure 1B), and we used it as a covariate of no-interest in neuroimaging Model 2. Using a one-way ANOVA with the Distance as the dependent variable and the Species as the independent variable, the main effect of Species was significant (F(3,88) = 15.84, p < 0.001). All between-species differences were significant (0.01 < p < 0.001; see Figure 1 and Supplementary file 1B).

These data are the topic of a publication about the impact of acoustic parameters on the recognition of the affective cues of primate vocalizations by human participants (Debracque et al., 2023). All details are described in this article for the present stimuli.

Most discriminant low-level acoustic parameters (neuroimaging Model 3)

Request a detailed protocol

Following the GDA on the 88 acoustic parameters (Eyben et al., 2015; Eyben et al., 2010) of the species stimuli presented above, we decided to use as covariates of no-interest the most discriminant low-level acoustic features of our stimuli to maximize brain activations that are independent of these features. We therefore included the most significant acoustic features ([r > 0.70] and [r < −0.70]) of the first three factors of the GDA, that explained 27.14%, 21.63% and 18.99% of the variance, respectively (Debracque et al., 2023). Such selection left us with the following acoustic features: Factor 1, (1) vocalization loudness (r = 0.92), (2) intensity (r = 0.87), (3) change in spectrum (r = 0.72); Factor 2, (4) bandwidth contour of the second formant (F2; r = 0.79); Factor 3, (5) power of the fundamental frequency (F0; r = 0.80), and finally (6) the difference in intensity contour (r = −0.71). The acoustic parameters were used as covariates of no-interest (N = 6) in that specific order—namely, from the highest to lowest factor saturation—in neuroimaging Model 3.

Again, all details of this analysis are described in detail in a dedicated article for the present stimuli (Debracque et al., 2023), and the values are reported in Supplementary file 1A.

Acoustic parameters characterizing chimpanzee and bonobo calls

Request a detailed protocol

We performed a principal component analysis (PCA) on the 88 GeMAPS (Eyben et al., 2015; Eyben et al., 2010) acoustic parameters by using only the stimuli recorded from chimpanzee and bonobo individuals, in order to constrain the extraction of acoustic features of interest to these two species only. We again included the most significant acoustic features and took care of multicollinearity (excluding acoustic parameters highly correlating together with [r > 0.70] and [r < −0.70]) before PCA. This analysis left us with 13 principal components (PCs) accounting in total for at least 80% of the variance. Variance explained specifically for each PC (and its name and loading) was: PC1 (F0 Range (semitones), 0.349), 12.47%; PC2 (Spectral Slope 500–1500 Hz Variability, +0.441), 9.50%; PC3 (F1 Frequency Variability, +0.357), 8.87%; PC4 (Unvoiced Segment Duration SD, +0.432), 7.38%; PC5 (HNR Variability, +0.339), 6.83%; PC6 (Spectral Slope 0–500 Hz (mean), +0.474), 6.08%; PC7 (F1 Bandwidth Variability, +0.414), 5.81%; PC8 (F0 Rising Slope Variability, –0.403), 5.22%; PC9 (F2 Amplitude Variability (relative to F0), +0.387), 5.09%; PC10 (Spectral Slope 0–500 Hz Variability, +0.358), 4.38%; PC11 (F3 Amplitude Variability (relative to F0), +0.520), 4.10%; PC12 (Jitter Variability, +0.407), 3.77%; PC13 (H1–A3 Amplitude Ratio Variability, +0.470), 3.67%. Cumulative variance among the 13 PC was 83.16%.

Imaging data acquisition

Species categorization task

Request a detailed protocol

Structural and functional brain imaging data were acquired by using a 3T scanner Siemens Trio, Erlangen, Germany with a 32-channel coil. A 3D GR\IR magnetization-prepared rapid acquisition gradient echo sequence was used to acquire high-resolution (0.35 × 0.35 × 0.7 mm3) T1-weighted structural images (TR = 2400 ms, TE = 2.29 ms). Functional images were acquired by using fast fMRI, with a multislice echo planar imaging sequence with 79 transversal slices in descending order, slice thickness 3 mm, TR = 650 ms, TE = 30 ms, field of view = 205 × 205 mm2, 64 × 64 matrix, flip angle = 50°, bandwidth 1562 Hz/Px. In total for this task, 636 functional volumes of 79 slices were acquired for each participant for a total of 50,244 slices per participant. For our whole sample of 23 participants, 14,628 volumes were acquired for a grand total of 1,155,612 slices.

TVAs localizer task

Request a detailed protocol

Structural and functional brain imaging data were acquired by using a 3T scanner Siemens Trio, Erlangen, Germany with a 32-channel coil. A magnetization-prepared rapid acquisition gradient echo sequence was used to acquire high-resolution (1 × 1 × 1 mm3) T1-weighted structural images (TR = 1900 ms, TE = 2.27 ms, TI = 900 ms). Functional images were acquired by using a multislice echo planar imaging sequence with 36 transversal slices in descending order, slice thickness 3.2 mm, TR = 2100 ms, TE = 30 ms, field of view = 205 × 205 mm2, 64 × 64 matrix, flip angle = 90°, bandwidth 1562 Hz/Px. In total for this task, 230 functional volumes of 36 slices were acquired for each participant for a total of 8280 slices per participant. For our sample of 98 participants, 22,540 volumes were acquired for a grand total of 811,440 slices. For the sample-specific data (N = 23), 5290 volumes were acquired for a grand total of 190,440 slices.

Whole-brain data analysis

Species categorization task analysis within the TVAs

Request a detailed protocol

Functional images were analyzed with Statistical Parametric Mapping software (SPM12, Wellcome Trust Centre for Neuroimaging, London, UK). Preprocessing steps included realignment to the first volume of the time series, slice timing, normalization into the Montreal Neurological Institute (MNI; Collins et al., 1994) space using the DARTEL toolbox (Ashburner, 2007) and spatial smoothing with an isotropic Gaussian filter of 8 mm full width at half maximum. To remove low-frequency components, we used a high-pass filter with a cutoff frequency of 1/128 Hz. Three general linear models were used to compute first-level statistics, in which each event was modeled by using a boxcar function and was convolved with the hemodynamic response function, time-locked to the onset of each stimulus. In Model 1, separate regressors were created for all trials of each species (species factor: human, chimpanzee, bonobo, macaque vocalizations) and two covariates of no-interest each (mean fundamental frequency and mean energy of each species) for a total of 12 regressors. Finally, six motion parameters were included as regressors of no interest to account for movement in the data and our design matrix therefore included a total of 18 columns plus the constant term. The species regressors were used to compute simple contrasts for each participant, leading to separate main effects of human, chimpanzee, bonobo, and macaque vocalizations. Covariates were set to zero in order to model them as no-interest regressors. In Model 2, separate regressors were created for all trials of each species (species factor: human, chimpanzee, bonobo, macaque vocalizations) and one covariate of no-interest for each species (acoustic distance for each species relative to human voice stimuli) for a total of eight regressors. Six motion parameters were included as regressors of no interest to account for movement in the data and our design matrix therefore included a total of 14 columns plus the constant term. The species regressors were used to compute simple contrasts for each participant, leading to separate main effects of human, chimpanzee, bonobo, and macaque vocalizations excluding acoustic distance (the covariate was set to zero in order to model it as ‘of no-interest’). In Model 3, separate regressors were created for all trials of each species (species factor: human, chimpanzee, bonobo, macaque vocalizations) and six covariates of no-interest each (vocalization loudness, intensity, change in spectrum, bandwidth contour of the second formant (F2), power of the fundamental frequency (F0), and finally the difference in intensity contour) for a total of 28 regressors. Six motion parameters were included as regressors of no-interest to account for movement in the data and our design matrix therefore included a total of 34 columns plus the constant term. The species regressors were used to compute simple contrasts for each participant, leading to separate main effects of human, chimpanzee, bonobo, and macaque vocalizations. Covariates were set to zero in order to model them as no-interest regressors. Finally, for Model 4 we used a model-based approach in which we modeled behavioral data as covariate for each trial as the main regressor of interest. Specifically, we computed a modeling of the behavioral data in R—see Species categorization task model above in the section on behavioral data analysis. On the whole sample (N = 23), we computed a mixed-effects logistic regression in which we predicted response accuracy as a function of the interaction between vocalization species and context of production. The six most discriminant acoustic parameters (those of Model 3) were also included as fixed effects, and we controlled for trial-wise reaction times (random effect). The obtained regression values for each trial and participant were then fitted on the whole sample and used as trial-level covariate in the present model, representing the probability of correctly categorizing human, chimpanzee, bonobo, and macaque stimuli in our task. The design matrix for each participant therefore included two regressors: (1) each trial onset and (2) the associated behavioral response fitted regression value. The advantage of this approach is that it includes all responses and all trials, including correct and incorrect responses; therefore, providing a clear and representative picture of the brain mechanisms underlying species categorization in our task—without losing power. Each model was independent from the others and was computed by rerunning the analysis based on each of the three design matrices, taking as data the final preprocessed data for each participant.

For each first-level model of analysis Models 1–3, each of their respective four simple contrasts were then taken to two flexible factorial second-level analyses. For all of these second-level analyses, there were two factors: Participants factor (independence set to yes, variance set to unequal) and the Species factor (independence set to no, variance set to unequal). For Model 4, the second-level modeling was a one-sample t-test of the parametric modulator, namely the fitted regression values of the probability of correctly categorizing the species stimuli—computed on the whole sample of 23 participants. Variance was also set to unequal in this model. For these analyses and to be consistent, we only included participants who were above chance level (25%) in the species categorization task (N = 23).

Brain region labeling was defined using xjView toolbox (http://www.alivelearn.net/xjview) implementing the Automated anatomical labeling (‘aal3’) atlas (Rolls et al., 2020). All neuroimaging activations were thresholded in SPM12 by using a whole brain (Models 1–3) or TVA inclusive masking (Model 4) voxelwise FDR correction at p < 0.05 and an arbitrary cluster extent of k>10 voxels to remove very small clusters of activity. For Models 1–3, the TVA outlined in the figures are therefore only visual outlines of these regions, but no ROI analysis was performed here in order to maximize data representativeness—which is impacted negatively by ROI analysis.

TVAs localizer task

Request a detailed protocol

Functional images were analyzed with Statistical Parametric Mapping software (SPM12, Wellcome Trust Centre for Neuroimaging, London, UK). Preprocessing steps included realignment to the first volume of the time series, slice timing, normalization into the MNI (Collins et al., 1994) space using the DARTEL toolbox (Ashburner, 2007) and spatial smoothing with an isotropic Gaussian filter of 8 mm full width at half maximum. To remove low-frequency components, we used a high-pass filter with a cutoff frequency of 1/128 Hz. A general linear model was used to compute first-level statistics, in which each block was modeled by using a block function and was convolved with the hemodynamic response function, time-locked to the onset of each block. Separate regressors were created for each condition (vocal and non-vocal; condition factor). Finally, six motion parameters were included as regressors of no interest to account for movement in the data. The condition regressors were used to compute simple contrasts for each participant, leading to a main effect of vocal and non-vocal at the first-level of analysis: [1 0] for vocal, [0 1] for non-vocal. These simple contrasts were then taken to a flexible factorial second-level analysis in which there were two factors: Participants factor (independence set to yes, variance set to unequal) and the Condition factor (independence set to no, variance set to unequal). An identical analysis architecture was used to delineate more specific subregions of the TVA according to the type of non-vocal material, with categories: animal sounds, music, nature sounds, artificial noise sounds. Timing onsets of these newly created ‘events’—including their duration—within each non-vocal block of the task were determined, and each main effect contrast computed at the first-level was then taken to a second-level flexible factorial analysis with settings identical to the above. All neuroimaging activations were thresholded in SPM12 by using a voxelwise FDR correction at p < 0.05 and an arbitrary cluster extent of k > 10 voxels to remove very small clusters of activity. Activation outline for vocal > non-vocal—the contrast revealing the TVA—was precisely delineated for the N = 98 and the N = 23 samples and overlaid on brain displays of the species categorization task (see Figure 5). TVA subregions according to the auditory material (specific non-vocal condition blocks material) are reported in Figure 5—figure supplement 1.

Figure 5 with 1 supplement see all
Temporal voice areas for the present study.

Enhanced whole-brain activity for voice compared to non-voice stimuli in the main sample of the study (A–C, N = 23), in an independent sample of N=98 participants (D–F) and the overlap between these two samples (G–I). Data corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05, k = 10 voxels minimum per cluster. TVA: temporal voice areas.

Data availability

All data, stimuli, and codes used in this article are available in the FAIR-compliant open repository YARETA (URL: https://doi.org/10/g8wkrb).

The following data sets were generated
    1. Ceravolo L
    2. Debracque C
    3. Grandjean D
    (2024) YARETA
    Data and code for study on anterior temporal voice areas' sensibility to chimpanzee calls.
    https://doi.org/10.26037/yareta:udkd6r7wlbeozey5rl7on4h6y4

References

    1. Belin P
    (2006) Voice processing in human and non-human primates
    Philosophical Transactions of the Royal Society B 361:2091–2107.
    https://doi.org/10.1098/rstb.2006.1933
    1. Collins DL
    2. Neelin P
    3. Peters TM
    4. Evans AC
    (1994)
    Automatic 3D intersubject registration of MR volumetric data in standardized Talairach space
    Journal of Computer Assisted Tomography 18:192–205.
  1. Software
    1. Lenth RV
    (2019)
    Emmeans: estimated marginal means, aka least-squares means, version 1.3.2
    R Package Version.
    1. Milham M
    2. Petkov C
    3. Belin P
    4. Ben Hamed S
    5. Evrard H
    6. Fair D
    7. Fox A
    8. Froudist-Walsh S
    9. Hayashi T
    10. Kastner S
    11. Klink C
    12. Majka P
    13. Mars R
    14. Messinger A
    15. Poirier C
    16. Schroeder C
    17. Shmuel A
    18. Silva AC
    19. Vanduffel W
    20. Van Essen DC
    21. Wang Z
    22. Roe AW
    23. Wilke M
    24. Xu T
    25. Aarabi MH
    26. Adolphs R
    27. Ahuja A
    28. Alvand A
    29. Amiez C
    30. Autio J
    31. Azadi R
    32. Baeg E
    33. Bai R
    34. Bao P
    35. Basso M
    36. Behel AK
    37. Bennett Y
    38. Bernhardt B
    39. Biswal B
    40. Boopathy S
    41. Boretius S
    42. Borra E
    43. Boshra R
    44. Buffalo E
    45. Cao L
    46. Cavanaugh J
    47. Celine A
    48. Chavez G
    49. Chen LM
    50. Chen X
    51. Cheng L
    52. Chouinard-Decorte F
    53. Clavagnier S
    54. Cléry J
    55. Colcombe SJ
    56. Conway B
    57. Cordeau M
    58. Coulon O
    59. Cui Y
    60. Dadarwal R
    61. Dahnke R
    62. Desrochers T
    63. Deying L
    64. Dougherty K
    65. Doyle H
    66. Drzewiecki CM
    67. Duyck M
    68. Arachchi WE
    69. Elorette C
    70. Essamlali A
    71. Evans A
    72. Fajardo A
    73. Figueroa H
    74. Franco A
    75. Freches G
    76. Frey S
    77. Friedrich P
    78. Fujimoto A
    79. Fukunaga M
    80. Gacoin M
    81. Gallardo G
    82. Gao L
    83. Gao Y
    84. Garside D
    85. Garza-Villarreal EA
    86. Gaudet-Trafit M
    87. Gerbella M
    88. Giavasis S
    89. Glen D
    90. Ribeiro Gomes AR
    91. Torrecilla SG
    92. Gozzi A
    93. Gulli R
    94. Haber S
    95. Hadj-Bouziane F
    96. Fujimoto SH
    97. Hawrylycz M
    98. He Q
    99. He Y
    100. Heuer K
    101. Hiba B
    102. Hoffstaedter F
    103. Hong S-J
    104. Hori Y
    105. Hou Y
    106. Howard A
    107. de la Iglesia-Vaya M
    108. Ikeda T
    109. Jankovic-Rapan L
    110. Jaramillo J
    111. Jedema HP
    112. Jin H
    113. Jiang M
    114. Jung B
    115. Kagan I
    116. Kahn I
    117. Kiar G
    118. Kikuchi Y
    119. Kilavik B
    120. Kimura N
    121. Klatzmann U
    122. Kwok SC
    123. Lai H-Y
    124. Lamberton F
    125. Lehman J
    126. Li P
    127. Li X
    128. Li X
    129. Liang Z
    130. Liston C
    131. Little R
    132. Liu C
    133. Liu N
    134. Liu X
    135. Liu X
    136. Lu H
    137. Loh KK
    138. Madan C
    139. Magrou L
    140. Margulies D
    141. Mathilda F
    142. Mejia S
    143. Meng Y
    144. Menon R
    145. Meunier D
    146. Mitchell AJ
    147. Mitchell A
    148. Murphy A
    149. Mvula T
    150. Ortiz-Rios M
    151. Ortuzar Martinez DE
    152. Pagani M
    153. Palomero-Gallagher N
    154. Pareek V
    155. Perkins P
    156. Ponce F
    157. Postans M
    158. Pouget P
    159. Qian M
    160. Ramirez J
    161. Raven E
    162. Restrepo I
    163. Rima S
    164. Rockland K
    165. Rodriguez NY
    166. Roger E
    167. Hortelano ER
    168. Rosa M
    169. Rossi A
    170. Rudebeck P
    171. Russ B
    172. Sakai T
    173. Saleem KS
    174. Sallet J
    175. Sawiak S
    176. Schaeffer D
    177. Schwiedrzik CM
    178. Seidlitz J
    179. Sein J
    180. Sharma J
    181. Shen K
    182. Sheng W
    183. Shi NS
    184. Shim WM
    185. Simone L
    186. Sirmpilatze N
    187. Sivan V
    188. Song X
    189. Tanenbaum A
    190. Tasserie J
    191. Taylor P
    192. Tian X
    193. Toro R
    194. Trambaiolli L
    195. Upright N
    196. Vezoli J
    197. Vickery S
    198. Villalon J
    199. Wang X
    200. Wang Y
    201. Weiss AR
    202. Wilson C
    203. Wong T-Y
    204. Woo C-W
    205. Wu B
    206. Xiao D
    207. Xu AG
    208. Xu D
    209. Xufeng Z
    210. Yacoub E
    211. Ye N
    212. Ying Z
    213. Yokoyama C
    214. Yu X
    215. Yue S
    216. Yuheng L
    217. Yumeng X
    218. Zaldivar D
    219. Zhang S
    220. Zhao Y
    221. Zuo Z
    (2022) Toward next-generation primate neuroscience: A collaboration-based strategic plan for integrative neuroimaging
    Neuron 110:16–20.
    https://doi.org/10.1016/j.neuron.2021.10.015
  2. Book
    1. Team, R
    (2015)
    RStudio: Integrated Development for R
    Boston, MA: Rstudio. Com.

Article and author information

Author details

  1. Leonardo Ceravolo

    University of Geneva, Geneva, Switzerland
    Contribution
    Conceptualization, Data curation, Software, Formal analysis, Validation, Investigation, Visualization, Methodology, Writing – original draft, Writing – review and editing
    Contributed equally with
    Coralie Debracque
    For correspondence
    Leonardo.Ceravolo@unige.ch
    Competing interests
    No competing interests declared
    ORCID icon "This ORCID iD identifies the author of this article:" 0000-0003-0638-3981
  2. Coralie Debracque

    University of Geneva, Geneva, Switzerland
    Contribution
    Conceptualization, Data curation, Validation, Investigation, Writing – original draft, Writing – review and editing
    Contributed equally with
    Leonardo Ceravolo
    Competing interests
    No competing interests declared
  3. Thibaud Gruber

    University of Geneva, Geneva, Switzerland
    Contribution
    Resources, Supervision, Funding acquisition, Validation, Project administration, Writing – review and editing
    Contributed equally with
    Didier Grandjean
    Competing interests
    No competing interests declared
  4. Didier Grandjean

    University of Geneva, Geneva, Switzerland
    Contribution
    Resources, Supervision, Funding acquisition, Validation, Project administration, Writing – review and editing
    Contributed equally with
    Thibaud Gruber
    Competing interests
    No competing interests declared

Funding

Swiss National Science Foundation (CR13I1_162720)

  • Thibaud Gruber
  • Didier Grandjean

National Centre of Competence in Research (NCCR) (51NF40-104897)

  • Didier Grandjean

Swiss National Science Foundation (PCEFP1_186832)

  • Thibaud Gruber

The funders had no role in study design, data collection, and interpretation, or the decision to submit the work for publication.

Acknowledgements

We thank the Swiss National Science Foundation (SNSF) for supporting this interdisciplinary project (grant CR13I1_162720/1 to DG-TG), the National Centre of Competence in Research (NCCR) (51NF40-104897 to DG) hosted by the Swiss Center for Affective Sciences, as well as the Fondation Ernst et Lucie Schmidheiny supporting CD. TG was additionally supported by a grant of the SNSF during the final editing of this article (grant PCEFP1_186832). We also thank Katie Slocombe and Zanna Clay for providing the nonhuman primates auditory stimuli and Daphne Bavelier for her advice on the design of the species categorization task. We would like also to acknowledge the staff of the Brain and Behavior Laboratory at the University of Geneva where all data were acquired.

Ethics

Participants gave written informed consent for their participation in accordance with ethical and data security guidelines of the University of Geneva. The study was approved by the Ethics Cantonal Commission for Research of the Canton of Geneva, Switzerland (CCER) and was conducted according to the Declaration of Helsinki.

Version history

  1. Sent for peer review:
  2. Preprint posted:
  3. Reviewed Preprint version 1:
  4. Reviewed Preprint version 2:
  5. Version of Record published:

Cite all versions

You can cite all versions using the DOI https://doi.org/10.7554/eLife.108795. This DOI represents all versions, and will always resolve to the latest one.

Copyright

© 2025, Ceravolo, Debracque et al.

This article is distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use and redistribution provided that the original author and source are credited.

Metrics

  • 2,137
    views
  • 118
    downloads
  • 1
    citation

Views, downloads and citations are aggregated across all versions of this paper published by eLife.

Citations by DOI

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Leonardo Ceravolo
  2. Coralie Debracque
  3. Thibaud Gruber
  4. Didier Grandjean
(2026)
Sensitivity of the human temporal voice areas to nonhuman primate vocalizations
eLife 14:RP108795.
https://doi.org/10.7554/eLife.108795.3

Share this article

https://doi.org/10.7554/eLife.108795