Animal vocalisations are often assigned to discrete call categories through expert annotation, yet it remains unclear how well such labels correspond to measured acoustic variation. This issue is relevant for rodent ultrasonic vocalisations (USVs), used in studies of communication, affective state, and behavioural phenotyping. Here, we present a reproducible bioacoustic framework combining interpretable acoustic descriptors, nonlinear dimensionality reduction, unsupervised clustering, statistical analysis, and feature-importance estimation. We analysed 801 manually inspected vocalisation observations obtained from four pups assigned to the neonatal-handling condition. Expert-defined distress and simple calls showed substantial overlap in low-dimensional representations, with negligible label-based separation. Independently, k-means clustering of the standardised acoustic feature space with k = 2 yielded measurable separation but negligible correspondence with manual labels. This pattern persisted in leave-one-pup-out sensitivity analyses, although the small number of biological subjects limits population-level generalisation. Maximum frequency, manually marked peak frequency, spectral slope, and duration contributed most strongly to reproducing cluster assignments. Acoustic properties also differed between recording phases, with post-NH observations showing higher automatically estimated peak and maximum frequencies and shorter durations than pre-NH observations. Overall, discrete call labels did not fully describe variation across measured acoustic dimensions, supporting data-driven representations as complementary tools to expert annotation.
Rat ultrasonic vocalisations, bioacoustics, manual annotation, acoustic descriptors, unsupervised clustering, acoustic variation