Sensitivity of the human temporal voice areas to nonhuman primate vocalizations

5 figures, 2 tables and 2 additional files

Figures

Timecourse and results of the species categorization task and the control task testing for potential attentional biases.

(A) Detail of the timecourse of four trials of the species categorization task in non-representative order, including waveform and spectrogram graphs for one example stimulus of each species. (B) Behavioral results of the task (N = 23) showing the probability (Probab.) of correctly categorizing (categ.) each species’ vocalization, using the six acoustic covariates from Model 3 and reaction times as variables. (C) Histograms of the acoustic Mahalanobis distance data of each species including mean (numbers represent exact mean value). (D) Control task and its design in an independent sample of 28 participants using each species’ vocalization as an exogenous cue preceding the target to be detected as fast as possible, namely a short sine wave tone (600 Hz). (E) Results of the control task (N = 28), showing no biasing of attention by any of the species stimuli: no species triggered attentional capture, yielding to no attentional advantage for target detection. For the results plots, violin plots illustrate distribution fit, error bars the standard error of the mean (SEM). Points represent individual values. ITI: intertrial interval; Resp.: response; Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. **p < 0.01, ***p < 0.001, n.s.: non-significant.

Figure 2 with 5 supplements
Whole-brain results with TVA outlines when contrasting the processing of chimpanzee to other species’ vocalizations with vocalization loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference as trial-level covariates of no-interest (Model 3).

Enhanced brain activity on a sagittal view with activity for [chimpanzee > human, bonobo, macaque] (dark blue to green), [chimpanzee > bonobo, macaque] (brown to red with light yellow outline), and [macaque > human, bonobo, chimpanzee] (red to yellow) vocalizations, with outlines of TVA for voice > animal sounds (A, B), voice > nature sounds (C, D), voice > music (E, F), and voice > noise (G, H). Brain activations are independent of the most discriminant low-level acoustic parameters of the stimuli set (Debracque et al., 2023). Data corrected for multiple comparisons using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. White outline: sample-specific temporal voice areas (TVA; N = 23); Dotted black outline: sample-specific TVA, per sound category; Blue outline: areas selective to chimpanzee calls. ‘a’ prefix: anterior; ‘m’ prefix: mid; ‘p’ prefix: posterior; STG: superior temporal gyrus; STS: superior temporal sulcus; L: left hemisphere; R: right hemisphere.

Figure 2—figure supplement 1
Whole-brain results when contrasting the processing of chimpanzee to other species’ vocalizations with mean fundamental frequency and energy as trial-level covariates of no-interest (Model 1).

(A–C) Enhanced brain activity for human and chimpanzee compared to bonobo and macaque vocalizations (purple to yellow) on a sagittal view, overlaid with activity specific to chimpanzee vocalizations (dark blue to green). (D) Percentage of signal change for each individual and relevant species according to the contrast in the left anterior superior temporal gyrus (aSTG1). Box plots represent mean value (black line) and the standard error of the mean with distribution fit. (E–G) Direct comparison between human and chimpanzee vocalizations (human > chimpanzee: dark red to yellow; chimpanzee > human: dark green to yellow) as well as between chimpanzee calls vs bonobo and macaque calls (chimpanzee > bonobo and macaque: brown to red) on a sagittal render. (H) Percentage of signal change in the anterior superior temporal gyrus (aSTG2) when contrasting chimpanzee to human vocalizations for each individual and relevant species according to the contrast with box plots representing mean value (black line) and the standard error of the mean with distribution fit. Brain activations are independent of low-level acoustic parameters for all species (mean fundamental frequency ‘F0’ and mean energy of vocalizations). Data corrected for multiple comparisons using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. Percentage of signal change extracted at cluster peak including 9 surrounding voxels, selecting among these the ones explaining at least 85% of the variance using singular value decomposition. Circles represent individual values, boxplot represents the mean and its standard error, and half-violin plots show data distribution. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: temporal voice areas of an independent sample of N = 98. ‘a’ prefix: anterior; ‘m’ prefix: mid; ‘p’ prefix: posterior; STG: superior temporal gyrus; STS: superior temporal sulcus; L: left hemisphere; R: right hemisphere.

Figure 2—figure supplement 2
Whole-brain results when contrasting the processing of chimpanzee to other species’ vocalizations with Mahalanobis acoustic distance as trial-level covariates of interest (Model 2).

(A–C) Enhanced brain activity for human and chimpanzee compared to bonobo and macaque vocalizations (purple to yellow) on a sagittal view, overlaid with activity specific to chimpanzee vocalizations (dark blue to green). (D) Percentage of signal change for each individual and relevant species according to the contrast in the left anterior superior temporal gyrus (aSTG6). Box plots represent mean value (black line) and the standard error of the mean with distribution fit. (E–G) Direct comparison between human and chimpanzee vocalizations (human > chimpanzee: dark red to yellow; chimpanzee > human: dark green to yellow) as well as between chimpanzee calls vs bonobo and macaque calls (chimpanzee > bonobo and macaque: brown to red) on a sagittal render. (H) Percentage of signal change in the anterior superior temporal gyrus (aSTG8) when contrasting chimpanzee to human vocalizations for each individual and relevant species according to the contrast with box plots representing mean value (black line) and the standard error of the mean with distribution fit. Brain activations are covarying with the acoustic distance of each stimulus for all species. Data corrected for multiple comparisons using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. Percentage of signal change extracted at cluster peak including nine surrounding voxels, selecting among these the ones explaining at least 85% of the variance using singular value decomposition. Circles represent individual values, boxplot represents the mean and its standard error, and half-violin plots show data distribution. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: temporal voice areas of an independent sample of N = 98. ‘a’ prefix: anterior; ‘m’ prefix: mid; ‘p’ prefix: posterior; STG: superior temporal gyrus; STS: superior temporal sulcus; L: left hemisphere; R: right hemisphere.

Figure 2—figure supplement 3
Whole-brain results when contrasting the processing of chimpanzee to other species’ vocalizations with vocalization loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference as trial-level covariates of no-interest (Model 3).

(A–C) Enhanced brain activity for human and chimpanzee compared to bonobo and macaque vocalizations (purple to yellow) on a sagittal view, overlaid with activity specific to chimpanzee vocalizations (dark blue to green). (D) Percentage of signal change for each individual and relevant species according to the contrast in the left anterior superior temporal gyrus (aSTG10). Box plots represent mean value (black line) and the standard error of the mean with distribution fit. (E–G) Direct comparison between human and chimpanzee vocalizations (human > chimpanzee: dark red to yellow; chimpanzee > human: dark green to yellow) as well as between chimpanzee calls vs bonobo and macaque calls (chimpanzee > bonobo and macaque: brown to red) on a sagittal render. (H) Percentage of signal change in the anterior superior temporal gyrus (aSTG12) when contrasting chimpanzee to human vocalizations and when contrasting chimpanzee to bonobo and macaque calls (aSTG13) for each individual and relevant species according to the contrast with box plots representing mean value (black line) and the standard error of the mean with distribution fit. Brain activations are independent of the most discriminant low-level acoustic parameters of the stimuli set. Data corrected for multiple comparisons using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. Percentage of signal change extracted at cluster peak including 9 surrounding voxels, selecting among these the ones explaining at least 85% of the variance using singular value decomposition. Circles represent individual values, boxplot represents the mean and its standard error, and half-violin plots show data distribution. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: temporal voice areas of an independent sample of N = 98. ‘a’ prefix: anterior; ‘m’ prefix: mid; ‘p’ prefix: posterior; STG: superior temporal gyrus; STS: superior temporal sulcus; L: left hemisphere; R: right hemisphere.

Figure 2—figure supplement 4
Whole-brain activations specific to macaque calls for Model 3.

(A, B) Enhanced whole-brain activity for macaque compared to human, chimpanzee, and bonobo vocalizations. (C, D) Enhanced whole-brain activity for macaque compared to chimpanzee and bonobo vocalizations. Data corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05, k = 10. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: sample-specific (N = 23) temporal voice areas; PT: planum temporale; mSTG: mid superior temporal gyrus.

Figure 2—figure supplement 5
Whole-brain additional activations for Model 3.

Enhanced whole-brain activity for (A, B) bonobo compared to human, (C, D) macaque compared to human, (E, F) macaque compared to bonobo, and (G, H) macaque compared to chimpanzee vocalizations. Data corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05, k = 10. TVA: sample-specific (N = 23) temporal voice areas; mSTG: mid superior temporal gyrus; mSTS: mid superior temporal sulcus; PT: planum temporale.

Synthesis of mid-to-anterior TVA clusters of activity recruited specifically by the processing of chimpanzee and macaque vocalizations (Models 1–3).

Anterior superior temporal gyrus (aSTG) and sulcus (aSTS) clusters recruited for the processing of chimpanzee calls as opposed to human voices, bonobo, macaque calls (pink: Model 1; purple: Model 2; blue: Model 3) in the general TVA (A, B, N = 98) as well as in the sample-specific TVA (C, D, N = 23). Macaque results are only significant for Model 3 (teal: macaque vs all other species). Model 1: mean of fundamental frequency and energy (covariates of no-interest, N = 2); Model 2: acoustic distance (covariate of no-interest, N = 1); Model 3: acoustic parameters that characterize low-level acoustics of our stimuli following a discriminant analysis (covariates of no-interest, N = 6). Data are all corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05 with t-values ranging from 5 to 10. Hum: human; Chimp: chimpanzee; Bon: bonobo; Mac: macaque. TVA: temporal voice areas. Prefix ‘a’: anterior; ‘m’: mid. L / R: left / right hemisphere.

Model-based correlates of the probability of correct species categorization, within sample-specific TVA (Model 4).

Correlates of the probability of correct species categorization computed using model-based analysis technique for all species, as illustrated on sagittal renders for all species, including human (A, B), then specifically for chimpanzee calls (C, D), bonobo calls (E, F), and macaque calls (G, H). These correlates were constrained to the bounds of the sample-specific TVA (N = 23, black outline) using an inclusive masking procedure with correction for multiple comparison using voxelwise false discovery rate (FDR) at a threshold of p < 0.05. The colorbars represent t-value statistics. TVA: temporal voice areas. Prefix ‘a’: anterior; ‘m’: mid; ‘p’: posterior. STG: superior temporal gyrus; STS: superior temporal sulcus; MTG: middle temporal gyrus; PT: planum temporale. L/R: left/right hemisphere.

Figure 5 with 1 supplement
Temporal voice areas for the present study.

Enhanced whole-brain activity for voice compared to non-voice stimuli in the main sample of the study (A–C, N = 23), in an independent sample of N=98 participants (D–F) and the overlap between these two samples (G–I). Data corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05, k = 10 voxels minimum per cluster. TVA: temporal voice areas.

Figure 5—figure supplement 1
Temporal voice areas locations and subregions as a function of the type of non-vocal material, sample-specific.

Enhanced whole-brain activity for voice compared to non-voice stimuli in the main sample of the study (N = 23), when non-vocal auditory material is animal sounds (A, B), music (C, D), nature sounds (E ,F), and artificial noises (G, H). Data corrected for multiple comparison using whole-brain voxelwise false discovery rate (FDR) at a threshold of p < 0.05. TVA: temporal voice areas.

Tables

Table 1
Post hoc contrasts for behavioral effects of the species by context interaction.
ContrastEstimateSEdfz ratiop-value(unc)p-value(FDR)
1bon ago–chimp ago–1.786446350.4294698Inf–4.159655763.187276e−056.010292e−05
2bon ago–hum ago–6.300502240.7182617Inf–8.771875121.757152e−182.899301e−17
3bon ago–mac ago–2.008753490.4126352Inf–4.868109971.126706e−062.478754e−06
4bon ago–bon ago–0.174143280.4128612Inf–0.421796196.731738e−017.052297e−01
5bon ago–chimp ago–1.983298860.4882581Inf–4.061988864.865640e−058.679249e−05
6bon ago–hum ago–5.735457800.7448880Inf–7.699758161.363240e−149.997093e−14
7bon ago–mac ago–2.371159910.3947683Inf–6.006460381.896173e−095.688518e−09
8bon ago–bon aff–0.504695070.4196264Inf–1.202724732.290829e−012.799902e−01
9bon ago–chimp aff–1.822681560.4239757Inf–4.299023581.715522e−053.431044e−05
10bon ago–hum aff–6.423269280.7143960Inf–8.991188562.445710e−199.704562e−18
11bon ago–mac aff–1.165027520.4263215Inf–2.732744226.280909e−039.421363e−03
12chimp ago–hum ago–4.514055900.6987376Inf–6.460301651.044945e−104.056844e−10
13chimp ago–mac ago–0.222307140.3099024Inf–0.717345684.731608e−015.384244e−01
14chimp ago–bon ago1.612303060.3982712Inf4.048253765.160118e−058.962311e−05
15chimp ago–chimp ago–0.196852510.2903578Inf–0.677965354.977937e−015.475730e−01
16chimp ago–hum ago–3.949011450.7019870Inf–5.625476651.849964e−084.883906e−08
17chimp ago–mac ago–0.584713570.3413709Inf–1.712839418.674209e−021.144996e−01
18chimp ago–bon aff1.281751280.4624629Inf2.771576665.578553e−038.562431e−03
19chimp ago–chimp aff–0.036235210.2699085Inf–0.134250008.932049e−019.069465e−01
20chimp ago–hum aff–4.636822930.6402440Inf–7.242274474.412218e−132.393656e−12
21chimp ago–mac aff0.621418820.3143750Inf1.976679844.807783e−026.475789e−02
22hum ago–mac ago4.291748750.7091699Inf6.051791631.432437e−094.727042e−09
23hum ago–bon ago6.126358960.6829140Inf8.970908892.940776e−199.704562e−18
24hum ago–chimp ago4.317203380.7105662Inf6.075722881.234304e−094.287583e−09
25hum ago–hum ago0.565044440.8198583Inf0.689197674.906989e−015.475730e−01
26hum ago–mac ago3.929342330.6815967Inf5.764907788.170250e−092.344507e−08
27hum ago–bon aff5.795807180.6925633Inf8.368631165.829140e−177.694464e−16
28hum ago–chimp aff4.477820680.6986217Inf6.409507101.459909e−105.352998e−10
29hum ago–hum aff–0.122767030.8810853Inf–0.139336158.891845e−019.069465e−01
30hum ago–mac aff5.135474720.6894527Inf7.448624929.431810e−145.659086e−13
31mac ago–bon ago1.834610210.4117633Inf4.455496838.369913e−061.781982e−05
32mac ago–chimp ago0.025454630.3645522Inf0.069824379.443334e−019.443334e−01
33mac ago–hum ago–3.726704310.7256292Inf–5.135824282.809100e−076.621451e−07
34mac ago–mac ago–0.362406420.2833893Inf–1.278828812.009573e−012.502488e−01
35mac ago–bon aff1.504058430.4754706Inf3.163304721.559890e−032.573818e−03
36mac ago–chimp aff0.186071930.3152714Inf0.590196065.550592e−016.005559e−01
37mac ago–hum aff–4.414515780.6599336Inf–6.689333192.241898e−119.247828e−11
38mac ago–mac aff0.843725970.3103779Inf2.718383076.560184e−039.621603e−03
39bon ago–chimp ago–1.809155580.4258084Inf–4.248755462.149614e−054.172780e−05
40bon ago–hum ago–5.561314520.6705264Inf–8.293953131.095456e−161.205002e−15
41bon ago–mac ago–2.197016630.3849704Inf–5.706975831.150011e−083.162530e−08
42bon ago–bon aff–0.330551780.3533299Inf–0.935533073.495136e−014.119268e−01
43bon ago–chimp aff–1.648538280.4027399Inf–4.093307944.252623e−057.796476e−05
44bon ago–hum aff–6.249125990.7090878Inf–8.812908981.219396e−182.682670e−17
45bon ago–mac aff–0.990884240.3686192Inf–2.688097417.186043e−031.031041e−02
46chimp ago–hum ago–3.752158940.6870573Inf–5.461201974.729216e−081.200493e−07
47chimp ago–mac ago–0.387861050.3913065Inf–0.991195013.215904e−013.859084e−01
48chimp ago–bon aff1.478603790.4980115Inf2.969015532.987555e−034.809235e−03
49chimp ago–chimp aff0.160617300.3028305Inf0.530386785.958438e−016.342853e−01
50chimp ago–hum aff–4.439970420.6631146Inf–6.695630862.147432e−119.247828e−11
51chimp ago–mac aff0.818271340.3244697Inf2.521873231.167318e−021.639212e−02
52hum ago–mac ago3.364297890.6872616Inf4.895221569.819503e−072.234784e−06
53hum ago–bon aff5.230762730.7002025Inf7.470357417.997730e−145.278501e−13
54hum ago–chimp aff3.912776240.7175861Inf5.452692314.961287e−081.212759e−07
55hum ago–hum aff–0.687811480.9013979Inf–0.763049764.454337e−015.157654e−01
56hum ago–mac aff4.570430280.6688791Inf6.832969048.317491e−123.921103e−11
57mac ago–bon aff1.866464850.4329373Inf4.311167061.623952e−053.349400e−05
58mac ago–chimp aff0.548478350.3478039Inf1.576975831.148011e−011.485661e−01
59mac ago–hum aff–4.052109360.6704754Inf–6.043636111.506791e−094.735630e−09
60mac ago–mac aff1.206132390.3120989Inf3.864583371.112790e−041.883183e−04
61bon aff–chimp aff–1.317986490.4580890Inf–2.877140724.012966e−036.306089e−03
62bon aff–hum aff–5.918574210.7345777Inf–8.057110747.811887e−167.365494e−15
63bon aff–mac aff–0.660332460.4377438Inf–1.508490661.314290e−011.668137e−01
64chimp aff–hum aff–4.600587720.6360309Inf–7.233276644.714776e−132.393656e−12
65chimp aff–mac aff0.657654040.3282083Inf2.003769994.509470e−026.200522e−02
66hum aff–mac aff5.258241750.6759408Inf7.779145937.301580e−156.023803e−14
  1. SE: standard error (of the mean); df: degrees of freedom; p-value(unc): uncorrected p-value; p-value(FDR): p-value corrected for multiple comparisons. hum: human; chimp: chimpanzee; bon: bonobo; mac: macaque; aff: affiliative context (positive for nonhuman primates; happy voices for human); ago: agonistic context (threat and distress for nonhuman primates; angry and fearful voices for human).

Table 2
Activations, cluster size, and coordinates for each contrast of interest of Model 3 (vocalization loudness, intensity, change in spectrum, F2 bandwidth contour, F0 power, and intensity contour difference as trial-level covariates of no-interest) in the sample-specific temporal voice areas, whole-brain voxelwise p < 0.05 FDR corrected, k > 10.
MNI coordinates
Region labelHemisphereXYZt-valueCluster size (voxels)
Chimpanzee > human, bonobo, macaque
Superior temporal gyrus ant10L–52–2–124.5472
Superior temporal gyrus ant11R540–123.1919
Chimpanzee > bonobo, macaque
Superior temporal gyrus ant13L–54–2–125.03100
Superior temporal gyrus antL–58–12–24.98
Superior temporal gyrus antL–522–124.38
Superior temporal gyrus ant14R58–2–124.9791
Superior temporal sulcus antR54–8–124.3
Chimpanzee > human
Superior temporal gyrus ant12L–50–4–122.710
Superior temporal gyrus antL–48–10–122.6
Human > chimpanzee
Superior temporal gyrus midL–54–1448.555392
Supramarginal gyrusL–60–48247.54
Superior temporal gyrus postL–54–58186.88
Middle temporal gyrus midL–66–22–86.39
Middle temporal gyrus postL–66–44–25.49
Supramarginal gyrusR56–42288.225250
Supramarginal gyrusR56–42367.52
Superior temporal gyrus midR56–827.5
Superior temporal gyrus postR50–48226.23
Superior temporal gyrus postR58–46146.22
Superior temporal gyrus midR50–1866.05
Superior temporal gyrus postR60–46185.99
Middle temporal gyrus postR70–3205.84
Middle temporal gyrus midR64–22–125.75
Middle temporal gyrus midR66–20–185.42
  1. ant: anterior; mid: central part; post: posterior.

  2. 10–14Figure S3 cluster labels.

Additional files

Supplementary file 1

Supplementary tables including acoustic details and additional MRI data brain coordinates.

Supplementary file 1 contains: Table A, acoustic parameters selected among the eGemaps corpus using a general discriminant analysis (N = 16); Table B, Mahalanobis distance estimates for the two-way interaction between species and social context factors; Table C, MNI cluster coordinates for contrasts of interest of Model 1—including mean of fundamental frequency and energy as trial-level covariates; Table D, MNI cluster coordinates for contrasts of interest of Model 2—including acoustic distance for the human voice as trial-level covariate; Table E, acoustic differences between chimpanzee and bonobo calls, per social context.

https://cdn.elifesciences.org/articles/108795/elife-108795-supp1-v1.docx
MDAR checklist
https://cdn.elifesciences.org/articles/108795/elife-108795-mdarchecklist1-v1.docx

Download links

A two-part list of links to download the article, or parts of the article, in various formats.

Downloads (link to download the article as PDF)

Open citations (links to open the citations from this article in various online reference manager services)

Cite this article (links to download the citations from this article in formats compatible with various reference manager tools)

  1. Leonardo Ceravolo
  2. Coralie Debracque
  3. Thibaud Gruber
  4. Didier Grandjean
(2026)
Sensitivity of the human temporal voice areas to nonhuman primate vocalizations
eLife 14:RP108795.
https://doi.org/10.7554/eLife.108795.3