Figures and data

Evolutionary context for speech emergence and acoustic analysis pipeline.
(A) Phylogenetic tree showing human autapomorphies relative to extant non-human primates. Shifts in the neural underpinnings of vocal control are marked in dark blue; peripheral shifts in vocal anatomy are marked in light gray. (B) Analysis pipeline for vocalization recordings. First (Preprocessing and segmentation panel), each recording was stripped of all possible extraneous noise and segmented into regions of quasi-continuous sound (top and bottom left). Then each region was split into a series of partially overlapping Hamming windows with a length of 0.02 s and an overlap of 50% (bottom right). Next (Acoustic analysis panel), each window was converted into 11 Mel-frequency cepstral coefficients. Principal component analysis (PCA) was then performed on this centered, scaled dataset of acoustic features. Finally (Volume estimation panel), the acoustic-feature space occupied by each vocalization type was estimated using convex hulls and probabilistic hypervolumes. PCs 1–5 were used to estimate these volumes. [Note to publisher: please print this figure in color.]

Recording sources.

Acoustic data used in analyses of acoustic feature space.
The table depicts the number of original recordings downloaded from each data source (“Recordings (n)”), the number of quasi-continuous vocal segments that each group of recordings was divided into following pre-processing (“Audio segments (n)”), and the number of partially overlapping Hamming windows (length = 0.02 s; overlap = 50%) that each group of audio segments was divided into (“Time windows (n)”).

Non-linguistic vocalizations occupy more acoustic-feature space than speech, while speech and non-human primate outgroup repertoires occupy similar amounts.
(A) Convex hull and probabilistic hypervolume volumes from a PCA including all data and constructed from MFCC 2–12. Non-linguistic sounds occupied more acoustic-feature space than speech, but the former group also contained more observations. Speech and chacma baboons occupied similar amounts of acoustic-feature space despite the former having more than 6x as many observations. (B) Differences in hypervolume size. Green color denotes significant differences; arrows point to the larger hypervolume in each dyad. Non-linguistic sounds occupied more acoustic-feature space than all other vocalization types in both analyses. No other significant group differences were recorded across both hypervolume and convex hull analyses. (C) Two-dimensional projections of hypervolumes for speech (orange) and non-linguistic vocalizations (blue). The larger size of the non-linguistic hypervolume is evident from the shape outlines. (D) Two-dimensional projections of hypervolumes for speech (orange) and chacma baboon vocalizations (green). The approximately equal size of the two hypervolumes is evident from the shape outlines. Chacma baboon was the non-human primate whose vocal repertoire was most fully captured in the dataset. [Note to publisher: please print this figure in color.]
