#B223360 # Abstract

In the HCRC Map Task Corpus dialogue recording between two speakers, a Male speaker B giving directions to a Female speaker A following directions. Both speakers, the giving and following, recorded separate word lists repeating words and phrases mentioned during the Map Task, without context. Each speaker repeated their list twice for more reliable measures to support several claims about differences in phonetic properties. In this report the speakers’ wave form and spectrograms are useful in understanding the relative Vowel space, Consonant Voice Onset Time(VOT), and Prosodic structural differences of each of the speakers. Especially my main discussion point… .

Annotation

The conventions followed while analysing both speakers’ word list readings and dialogue, allowed a simple, replicable, algorithm to decide the time of intervals and point’s in the audio. For every boundary in the waveform, choosing an interval of some number of complete cycles (ex. peak to peak, zero-crossing to zero-crossing) is conventionally the zero-crossing on the rise of the wave for every point in this report. As well as for every segmentation of a phone or syllable(which may include multiple phones, and are transcribed with each corresponding phonetic symbol), the segment will consistently begin and end at the point in the rising waveform as it passes through zero from negative to positive. For my monophthongs, I found most meaningful results from using the conventional midpoint in time per vowel. When vowels are presented in a word that includes voiced “r” or “l” coloring, finding the relative valley in the waveform’s amplitudes between the vowel and consonant/approximate coloring.

Vowels

Both speakers utilise a wide range of vowel spaces to filter their words. This rearrangement of the filter, as more clearly understood by the IPA vowel chart, allows the speakers to distinguish their diction during both repetitions of their respecive word lists. Filtering is understood by the Source-Filter Theory, by blah blah, of changing the shape of the tongue lips and soft palate which manipulates the wave of the sound passing through our vocal tract. This is useful in my approach to the vowel space, while following my conventions of locating the start and end of monopthongs, I used the midpoint between timstamps to plot each speaker’s vowel spaces.

To better understand the vowel space of each speaker, it is vital to pay attention to the formants created on the spectrogram from dense areas of high pressure. For the purposes of these cardinal- and some mid vowels, the first two formants are suffient to compare. In each repetition blah blah blah. Compared to speaker B, blah blah blah.

Then to visualize the differences among the speaker’s vowel spaces, we plotted them in comparison.

1. “Chapel”

“The vowel monophthong (a) is consistently identified by ‘chapel’ repeated twice by both speaker A and B. The first vowel in the word is an open, front, and un-rounded vowel. We identified the formants 1 and 2 and averaged from both repetitions their frequencies for both speakers to plot the vowel space. Speaker A’s a average frequencies at the midpoint are F1=(819.4Hz) and F2=(1596Hz), and speaker B’s a averages at the midpoint are F1=(689.6Hz) and F2=(1476Hz).

2. “Footbridge”

This figure of the word, “footbridge” includes another monophtong in the first syllable both speakers spoke the secondary cardinal vowel #10(\(ʊ\)), a close-mid-closed(speaker A’s second repetiton central vowel #17(\(\istrike\))), front, and rounded vowel. For speaker A, the average first and second formants’ frequencies at the midpoint F1=(490.1Hz) and F2=(2013.5Hz). And for speaker B’s \(ʊ\) average F1=(344.4Hz) and F2=(1552.5Hz).

3. “Farmyard”

The first syllable of “Farmyard” identifies an unrounded version of the first monophthong spoken in relative continuity(ranging from a nearly-open central vowel to an open unrounded back vowel between both speakers as opposed to the second syllable’s rounded back vowel. The back, rounded, and open monophthong considered here is denoted as the secondary cardinal vowel #13(\(\alpha\)). Identified in this word by the rounding during the first syllable, “farm” transition to the back, unrounded and open second vowel(primary cardinal vowel #5(\(ɒ\)) in, “yard”. Interestingly, both speakers produce a more central variation during their first syllable. This may be according to the speakers’ scottish dialect. The average first and second formants’ frequencies at the midpoint are F1=(573.5Hz) and F2=(1137Hz) for speaker A. And for speaker B average F1=(497.9Hz) and F2=(1265Hz).

4. “Cattle Stockade”

The first syllable of the second word in “Cattle Stockade” is a long, back, rounded, and low-mid vowel spoken consistently among the two speakers. Identified by primary cardinal vowel #6 (\(ø\)), both speakers emulate a transition from the front vowel shape in the first syllable of “cattle” to the back vowel shape in the first syllable of “stockade”. To plot the formants of this monopthong, the averages at the midpoint are found. For speaker A’s \(ø\) average F1=(459.5Hz) and F2=(1092.5Hz), speaker B’s \(ø\) average F1=(555.4Hz) and F2=(1035Hz).

5. & 6. “Indian Country”

The figure above transcribes “Indian Country” including two consistent monophthongs among each instance of both speakers. In the first syllable of “Indian” is a distinctive sound in the English language producing an unrounded, mostly closed, front vowel (\(ɪ\)). According to the spectrogram for speaker A, the average first and second formants’ frequencies at the midpoint of \(ɪ\) are F1=(536.2Hz) and F2=(2339Hz). And for speaker B average F1=(421.1Hz) and F2=(1571.5Hz) at the midpoint of \(ɪ\).

In the first syllable of the second word, “Country” is another monopthong that presents itself. Denoted as primary cardinal vowel #14(\(ʌ\)), both speakers iterate an unrounded, close-mid For speaker A the average \(ʌ\) F1=(517Hz) and F2=(1380Hz), and speaker B’s averages at the \(ʌ\) midpoint are F1=(555.35Hz) and F2=(1207.5Hz).

7. “Abandoned Truck”

“Abandoned Truck” exhibits another consistently produced vowel among both speakers in the very first phonetic utterance in the phrase (\(ə\)). The ‘schwa’ is the most central and generally distinguished vowel, that is neither front nor back, neither entirely open nor closed, and usually unrounded (when spoken in English). To plot the first two formants of both speakers, speaker A’s average frequencies at the midpoint are F1=(1009.95Hz) and F2=(1821.5Hz) and speaker B’s average F1=(651.25Hz) and F2=(1802Hz).

8. “Baboons”

“Baboons” highlights the primary cardinal vowel #8 in the IPA, a long u in the second syllable. Spoken with a relatively consistent closure, the monophthong is a closed rounded back vowel. The average frequencies of u at the midpoint are F1=(459.5Hz) and F2=(1802Hz) for speaker A, and F1=(344.4Hz) and F2=(1629Hz) for speaker B.

9. “Ravine”

The second syllable of the word “Ravine” highlights the first primary cardinal vowel in the IPA, a long i in the second syllable. Spoken with a relatively consistently unimpeded opening since it emulates the most open front cardinal vowel and described as an unrounded, front, open monophthong. According to the midpoints of i on the spectrogram, the formants for both speakers are as follows: For speaker A average F1=(440.3Hz) and F2=(2876Hz), and for speaker B average F1=(307.6Hz) and F2=(2087.5Hz).

10. “Seven Beeches”

Found in both the first syllable of “Seven” and the second syllable of “Beeches”, is consistent evidence of both speakers producing primary cardinal vowel #3 (\(ɛ\)). But after further comparison of the vowel spaces between this vowel and the second syllable of the second word “Beeches”, may be a more open, and accurate example.

For speaker A’s \(ɛ\) the first and second formants’ frequencies at the midpoint are on average F1=(460.7Hz) and F2=(2221.5Hz). And for speaker B’s \(ɛ\) average frequencies F1=(346.9Hz) and F2=(1838.5Hz).

##    Formant 1 Formant 2 Speaker       Word
## 1     819.40    1596.0       A     Chapel
## 2     490.10    2013.5       A Footbridge
## 3     573.50    1137.0       A   Farmyard
## 4     459.50    1092.5       A   Stockade
## 5     536.20    2339.0       A     Indian
## 6     517.00    1380.0       A    Country
## 7     651.25    1802.0       A  Abandoned
## 8     459.50    1802.0       A    Baboons
## 9     440.30    2876.0       A     Ravine
## 10    460.70    2221.5       A    Beeches
## 11    689.60    1476.0       B     Chapel
## 12    344.40    1552.5       B Footbridge
## 13    497.85    1265.0       B   Farmyard
## 14    555.40    1035.0       B   Stockade
## 15    421.10    1571.5       B     Indian
## 16    555.35    1207.5       B    Country
## 17   1169.00    2320.0       B  Abandoned
## 18    344.40    1629.0       B    Baboons
## 19    307.60    2087.5       B     Ravine
## 20    346.90    1838.5       B    Beeches

We can compare these speakers by comparing their respective vowel spaces for each vowel, (a,\(ʊ\),\(\alpha\),\(ø\),\(ɪ\),\(ʌ\),,u,\(ə\),i,\(ɛ\)) labeled by the example word.

For the cases of the schwa, the second speaker’s averages were higher than we might expect, and therefore are to be considered an outlier for this analysis. The x and y values were reversed to relocate the origin(0,0) to the upper left-hand corner, which is easier to compare the average vowel spaces to the IPA vowel chart quadrilateral, although to more clearly compare vowel spaces, an appropriate portion of the possible ranges of sound frequency are zoomed in. In this case of two Scottish speakers, we can identify a reasonable range of understanding per vowel, according to its openness(F1 higher values), front to back(F2 high to low values), and lip rounding as distinguished relatively by the IPA vowel pairs of symbols. Overall, the female speaker A speaks at generally higher pitches compared to the male speaker B which is congruent with the usual high-voice associated with women and vice-versa for men. In comparison to B’s vowel space, speaker A more often tended towards more central and closed vowels which may explain extreme cases, for example “Stockade” and “Country”, where the average distance in vowel space purely according to pitch is no longer sufficient to override the degree of closure. Other than what might be explained by the difference in “thickness” of each speaker’s scottish dialect, the speakers had a repetitive ratio between their respective F2 Frequency, exemplifying they adhere to relative borders between sounds that have a distinctive meaning. ## Consonant Voice Onset Time (VOT)

1. p

First, consider the phrase “totem pole” in the word list:

One key consonant to identify in this word is a p in the first sound of the second word. This plosive is voiceless and made bilabial, therefore the burst precedes the voicing. The VOT for speaker A p averages _s difference between the burst from the voicing time, and speaker B p VOT averages _s. # 1A = , 1B, 2A=,2B=. ### 2. b Take a look at the first sound of the second word in the phrase, “moored boats”.

##    bilabials      vot speaker count repetition bl_factor bl_speaker_factor
## 1       p_1A  0.08190       A     1          1      p_1A                 A
## 2       p_2A  0.10950       A     2          2      p_2A                 A
## 3       p_1B  0.04500       A     3          1      p_1B                 A
## 4       p_2B  0.04110       A     4          2      p_2B                 A
## 5    p A Avg  0.09570       A     5    Average   p A Avg                 A
## 6    p B Avg  0.04305       B     6    Average   p B Avg                 B
## 7       b_1A -0.03500       B     7          1      b_1A                 B
## 8       b_2A -0.04340       B     8          2      b_2A                 B
## 9       b_1B -0.05380       B     9          1      b_1B                 B
## 10      b_2B -0.02700       B    10          2      b_2B                 B
## 11   b A Avg -0.03920       A    11    Average   b A Avg                 A
## 12   b B Avg -0.04040       B    12    Average   b B Avg                 B
##    bl_countfactor bl_rep_factor
## 1               1             1
## 2               2             2
## 3               3             1
## 4               4             2
## 5               5       Average
## 6               6       Average
## 7               7             1
## 8               8             2
## 9               9             1
## 10             10             2
## 11             11       Average
## 12             12       Average

These first two consonants share is their place of articulation as bilabial plosives, however the two speakers may still have individual differences in VOT.This is a voiced version of the first bilabial p, (b). The VOT for speaker A p averages 0.0957s difference between the burst from the voicing time, and speaker B p VOT averages 0.0431s. The difference between the speakers’ voiceless bilabial onset time is 0.0527s.While the VOT for speaker A b averages -0.0392s difference between the burst from the voicing time, and speaker B b VOT averages -0.0404s. The difference between the speakers’ voiced bilabial onset time is -0.0012s.

3. t

4. d

##    alveolar_vot speaker count repetition av_speaker_factor av_countfactor
## 1       0.10210       A     1          1                 A              1
## 2       0.08540       A     2          2                 A              2
## 3       0.08920       A     3          1                 A              3
## 4       0.05240       A     4          2                 A              4
## 5       0.09375       A     5    Average                 A              5
## 6       0.07080       B     6    Average                 B              6
## 7      -0.01290       B     7          1                 B              7
## 8      -0.02240       B     8          2                 B              8
## 9      -0.02760       B     9          1                 B              9
## 10     -0.01850       B    10          2                 B             10
## 11     -0.01765       A    11    Average                 A             11
## 12     -0.02305       B    12    Average                 B             12
##    av_rep_factor
## 1              1
## 2              2
## 3              1
## 4              2
## 5        Average
## 6        Average
## 7              1
## 8              2
## 9              1
## 10             2
## 11       Average
## 12       Average

Alveolar place of articulation for these two plosives, however the two speakers may still have individual differences in VOT. This is a voiced version of the first bilabial t, (d). The VOT for speaker A t averages 0.0938s difference between the burst from the voicing time, and speaker B t VOT averages 0.0708s. The difference between each speakers’ voiceless bilabial onset time is 0.023s. While the VOT for speaker A d averages -0.0177s difference between the burst from the voicing time, and speaker B d VOT averages -0.0231s. The difference between the speakers’ voiced bilabial onset time is -0.0054s.

5. k

See figure 3.1, has another consistent example of a voiceless velar plosive (k) in the first sound of the second word “cottage”.

6. g

##    velar_vot speaker count repetition av_speaker_factor av_countfactor
## 1     0.0740       A     1          1                 A              1
## 2     0.0796       A     2          2                 A              2
## 3     0.0712       A     3          1                 A              3
## 4     0.0628       A     4          2                 A              4
## 5     0.0768       A     5    Average                 A              5
## 6     0.0670       B     6    Average                 B              6
## 7    -0.0258       B     7          1                 B              7
## 8    -0.0386       B     8          2                 B              8
## 9    -0.0396       B     9          1                 B              9
## 10   -0.0246       B    10          2                 B             10
## 11   -0.0322       A    11    Average                 A             11
## 12   -0.0321       B    12    Average                 B             12
##    av_rep_factor
## 1              1
## 2              2
## 3              1
## 4              2
## 5        Average
## 6        Average
## 7              1
## 8              2
## 9              1
## 10             2
## 11       Average
## 12       Average

Velar place of articulation for these two plosives, however the two speakers may still have individual differences in VOT. This is a voiced version of the first bilabial k, (g). The VOT for speaker A k averages 0.0768s difference between the burst from the voicing time, and speaker B k VOT averages 0.0670s. The difference between each speakers’ voiceless bilabial onset time is 0.0098s. While the VOT for speaker A g averages -0.0322s difference between the burst from the voicing time, and speaker B g VOT averages -0.0321s. The difference between the speakers’ voiced bilabial onset time is insignificantly small(<0.004s).

Prosodic Structure

Between the word lists and the dialogue, the way certain monophthongs and other sounds are iterated may be affected by the context of speaking.

Poisoned Stream

The vot for the “d” from the word list to the dialogue goes almost entirely voiceless, to the point that there is little to no break between the bursting and the subsequent fricative in the dialogue. (hypothesis) ### Remote Village

Abandoned Truck

Word list, compared to the dialogue:

Compare/contrast vot, vowel space, speakers. I wish I had time to finish the prosodic structure section, apologies.