Sonic Human-Robot Interaction aims at equipping robots with the ability to convey emotions and intentions via sounds. Such sounds are typically handcrafted by human experts, which results in expensive, small sound sets with limited expressivity. To overcome these limitations, in this paper we propose (i) a module for the analysis of the emotional content of robotic sounds (RER), which combines a Contrastive Learning model with a classifier and regressor, and (ii) an Evolutionary Strategy (ES) that enables the automated modulation of a sound to control for its emotional content. Experiments performed on 118 emotion-labelled sounds provided by a publicly available robot vocal library suggest that: (i) state-of-the-art solutions for emotion recognition in speech or musical samples cannot be directly applied to robotic sounds; (ii) our proposed RER module achieves an F1-score of 88 % in the categorical emotion classification of robotic sounds and correctly places sounds on the Valence-Arousal space; (iii) our proposed ES is a promising first step towards modulating a sound to convey a target emotion.