11.9.2 Synthesising vowels

Sound 1 in Section 11.9 gave some examples of synthesised vowel sounds, using a source-filter model. This section gives the details of what was done. First, the source. The waveform of volume flow rate through the vocal folds was approximated using a formula suggested by Titze [1]:

$$u(t) = \max[0, e^{\alpha t} (\beta + \sin{\omega t})] \tag{1}$$

where $\omega$ is the desired frequency and $\alpha$ and $\beta$ are constants that can be adjusted to produce different values of the open quotient and different degrees of skew in the pulse shapes. For the example used here, based on mechanism M1, the values $\alpha = 700$ and $\beta = 0.03$ were used. This combination gives an open quotient around 0.5 and a small degree of skew, as shown in Fig. 1, reproduced from Fig. 4 of Section 11.9.

Figure 1. Idealised plot of the volume flow rate past the vocal folds generated by eq. (1), reproduced from Fig. 4 of Section 11.9.

Three notes were synthesised, based on equal-tempered frequencies $A_2$, $B_2$ and $C \sharp_3$. However, in the interests of realism the waveforms for each note were not exactly periodic: a precisely periodic waveform gives an immediately “electronic” feel to the sound. Measurements by Ternström and Friberg [2] have shown that when a singer is asked to sing a steady note without vibrato, the sound they produce contains small fluctuations in the fundamental period. These fluctuations are not random, they contain structure: the frequency spectrum of the period lengths shows a peak in the vicinity of 5 Hz, with a bandwidth of a few Hz.

They recommended modelling the effect with a set of periods determined by filtering random noise with a first-order resonance. For the examples used here, the peak frequency of the first-order resonance was set at 6.5 Hz, and its Q-factor was set to 3. A set of random numbers scaled to have zero mean was filtered by an IIR filter with these parameters, then scaled to have unit standard deviation. The successive periods of the synthesised waveform were then given by scaling the nominal equal-tempered period by a factor $(1+\lambda)$ where $\lambda$ is the next number from the resulting list of filtered random numbers, multiplied by 0.003.

For each of the three notes, this procedure was followed to create a waveform of length 1 s. To avoid annoying clicks, the first 5 periods were linearly ramped in amplitude, starting from zero. The final 5 periods were similarly ramped back down to zero. The three notes were separated by a brief interval of silence.

This input waveform was then filtered by a linear IIR filter to represent each desired vowel configuration. The first three formants for each vowel were included, using frequencies taken from Ladefoged and Johnson [3], and Q-factors inspired by the work of Hanna, Smith and Wolfe [4]. Suitable modal amplitudes were determined empirically to achieve a satisfactory sound — there are no measured values available. The values of these parameters are all listed in Table 1.

“hard”“food”“bed”
$F_1$ (Hz)710310550
$Q_1$555
$A_1$211
$F_2$ (Hz)11008701770
$Q_2$101010
$A_2$-1.5-1.5-1.5
$F_3$ (Hz)254022502490
$Q_3$101010
$A_3$0.50.50.5
Table 1. Formant frequencies $F_j$, Q-factors $Q_j$ and modal amplitudes $A_j$ for formants $j = 1,2,3$, for the three vowels in the synthesis example.

[1] Ingo R. Titze: “Sensitivity of odd-harmonic amplitudes to open quotient and skewing quotient in glottal airflow (L)”; Journal of the Acoustical Society of America 137, 502–504 (2015).

[2] Sten Ternström and Anders Friberg; “Analysis and simulation of small variations in the fundamental frequency of sustained vowels”, Quarterly Progress and Status Reports 30, 3, 1—14, Department of Speech, Music and Hearing, KTH Stockholm (1989), available from http://www.speech.kth.se/qpsr

[3] P. Ladefoged and K. Johnson. “A course in phonetics”. Wadsworth (2011).

[4] Noel Hanna, John Smith and Joe Wolfe; “Low frequency response of the vocal tract: acoustic and mechanical resonances and their losses”, Proceedings of Acoustics 2012, Australian Acoustical Society (2012)