Measurement and comparison of acoustic space use in vocalizations of humans and close primate relatives

  1. Department of Integrative Biology; University of Texas at Austin, Austin, United States
  2. Smithsonian Tropical Research Institute, Balboa, Panama
  3. Department of Geological Sciences; University of Texas at Austin, Austin, United States

Peer review process

Not revised: This Reviewed Preprint includes the authors’ original preprint (without revision), an eLife assessment, and public reviews.

Read more about eLife’s peer review process.

Editors

  • Reviewing Editor
    Daniel Takahashi
    Universidade Federal do Rio Grande do Norte, Natal, Brazil
  • Senior Editor
    Andrew King
    University of Oxford, Oxford, United Kingdom

Reviewer #1 (Public review):

Summary:

The authors conducted a comparative acoustic analysis of primate vocal repertoires, focusing on the assumption that speech and language evolution required and involved an expansion in the acoustic space of voiced vocalizations from non-human primates to humans. Results challenge this idea. The study compiles and analyzes a large dataset of calls to quantify differences in vocal production space.

Strengths:

The study is technically sound, with a solid implementation of acoustic measurements and a valuable new dataset that brings empirical rigor to test a dominant, yet hitherto strictly theoretical, notion about what speech and language evolution entailed. It provides concrete comparative acoustic data across species to disprove that speech and language required an increase in the range of voiced calls, and thus, by extension, of vowels. The approach is methodologically rigorous and directly engages with the relevant data, rather than relying on untested presumptions of what great apes "ought" to be able to do or not.

Weaknesses:

The theoretical contextualization should be strengthened and updated, as several aspects contain inaccuracies, most notably by equating voiced calls or vocalizations with speech (overlooking the critical role of consonants, as human languages typically show vowel:consonant ratios of 1:4 or greater) and misrepresenting the premises and current status of the neural (Kuypers-Jürgens) hypothesis.

The discussion drifts into speculative territory on features like syntax and co-articulation that fall outside the paper's scope and data, and it does not sufficiently engage recent evidence on vocal learning and consonant-like capacities in great apes.

Minor issues include incomplete sampling justifications, imprecise terminology, and reliance on references that have been critiqued in more recent work.

Reviewer #2 (Public review):

This study examines the evolutionary context of the emergence of human speech. The authors address the widely held hypothesis that the expansion of the human vocal space, resulting from modifications of the vocal tract, was a key prerequisite for the evolution of spoken language.

To test this hypothesis, the authors quantified the acoustic space of human speech, non-linguistic vocalizations, and musical vocalizations and compared it with that of nonhuman primates, chimpanzees, bonobos, and chacma baboons.

The authors found that speech and song occupied significantly less volume in the acoustic space than human non-linguistic vocalizations. In addition, the acoustic-feature volume of speech and song was not statistically distinct from that of non-human primates. Accordingly, the authors conclude that the evolution of human speech did not depend on an expansion of the human vocal acoustic space.

I find the analysis presented in this manuscript highly convincing. It is conducted at a contemporary scientific standard, and the results provide strong support for the authors' conclusions. I particularly appreciate that the authors explicitly discuss the limitations of their approach. For example, they acknowledge that MFCCs cannot capture all aspects of acoustic structure.

I have only three minor comments:

First, the authors may wish to briefly summarize the main findings of the study by Anikin et al., as it represents the central reference for the present work. A concise summary in two or three sentences would help readers who are not familiar with that study.

Second, I would appreciate a brief explanation of why the authors chose this particular statistical approach.

Third, the authors could briefly mention that the Chacma baboon dataset provides a very comprehensive representation of the vocal repertoire of this species, although a small number of rare vocalizations are not included. I am not sure whether a similar limitation also applies to the chimpanzee and bonobo datasets, but if so, it would be useful to mention this as well.

  1. Howard Hughes Medical Institute
  2. Wellcome Trust
  3. Max-Planck-Gesellschaft
  4. Knut and Alice Wallenberg Foundation