Добавил:
kiopkiopkiop18@yandex.ru t.me/Prokururor I Вовсе не секретарь, но почту проверяю Опубликованный материал нарушает ваши авторские права? Сообщите нам.
Вуз: Предмет: Файл:
Ординатура / Хирургия / Библиотека им академика М.И. Перельмана / Книга_4506_Библиотеки_им_академика_М_И_Перельмана.pdf
Скачиваний:
0
Добавлен:
31.08.2026
Размер:
34 Мб
Скачать
176 The Voice and Voice Therapy
FIGURE 68. Narrowband spectrogram of a voice with normal quality. Used
with the permission of Pentax Medical.
blends time changes, making it great for seeing individual harmonics. Wideband filtering, on the other hand, provides a clearer view of timing but merges frequencies together. For studying voice disorders, narrowband spectrograms are useful because they reveal how steadily the vocal folds are vibrating by showing changes in the voice’s harmonic structure. Figure 6–9 shows a spectrogram of an 18-year-old female with bilateral vocal fold nodules. Figure 6–10 shows a spectrogram of a 70-year-old male with unilateral vocal fold paralysis. Note that in both Figure 6–9 and 6–10, the harmonics are not clear or strong, and there is “noise” between the harmonics. The speaker and listener would likely perceive these voices as having an abnormal vocal quality.
The document “Recommended Protocols for Instrumental Assessment of Voice” by the ASHA Expert Panel (Patel et al., 2018) outlines protocols for the instrumental assessment of voice production. Its purpose is to recommend standardized procedures in laryngeal endoscopic imaging, acoustic analyses, and aerodynamic assessments. This standardization aims to improve the evidence for voice assessment measures, enable valid comparisons of results within and across clients and facilities, and facilitate the evaluation of treatment efficacy. The process combined existing evidence with expert consensus, informed by a survey of clinicians and peer review.
The section on Acoustic Analysis within the recommended protocols for instrumental assessment of voice production focuses on a systematic approach to capturing and analyzing the acoustic features of voice. This section focuses on data collection tasks, technical specifications for recording, and data analysis guidelines. We now describe each of these as well as additional measures with clinical utility.
To ensure the accuracy of acoustic measurements, specific technical requirements for recording equipment are outlined:
CHAPTER 6 Evaluation of the Voice 17 7
FIGURE 69. Narrowband spectrogram of a voice with vocal nodules. Used with the permission
of Pentax Medical.
FIGUR E 610. Narrowband spectrogram of a voice with vocal fold paralysis. Used with the
permission of Pentax Medical.
Microphone: A head-mounted omnidirectional microphone is recommended,
positioned 4 to 10 cm from the lips, to capture a clear and consistent acoustic signal.
Preamplifier and digital recording: The preamplifier should match the microphone’s
specifications, and digital recordings should have a minimum sampling rate of 44.1 kHz and a resolution of 16 bits or higher, ensuring high fidelity.
The Visi-Pitch (Model 3950c) with Computerized Speech Lab (CSL Model 4500b) (PENTAX
Medical Corp., Montvale, New Jersey) (Figure 6–12) is a clinical instrument widely used for
178 The Voice and Voice Therapy
FIGURE 611. Voice range profile. Used with the permission of
Pentax Medical.
FIGURE 612. A Visi-Pitch IV in clinical use. Used with the permission of
Pentax Medical.
CHAPTER 6 Evaluation of the Voice 17 9
voice assessment and treatment. It displays visual feedback of the voice’s pitch and intensity in real time, helping clients adjust their vocal production during therapy. The software analyzes and displays the range of pitch and intensity, offering insights into the vocal capabilities and limitations of the user. It can generate spectrograms, providing a visual representation of the frequency spectrum over time. This is valuable for examining the harmonic structure of the voice and identifying any irregularities. It includes spectral measures, which are important for assessing the quality of the voice and the presence of dysphonia. The Visi-Pitch is used in clinical settings to assess voice disorders, monitor changes in vocal function over time, and guide therapy. It is also used in research to study voice production and vocal health. The tool’s ability to provide immediate visual, auditory, and quantitative feedback makes it an effective aid in voice therapy, allowing for targeted interventions and helping clients visualize their progress. Monitoring important speech/voice behaviors with concrete visual displays helps clients reach therapy goals more easily.
The protocols specify methods for analyzing the collected acoustic data, focusing on key
measures of vocal function:
Habitual vocal SPL (sound pressure level): Vocal intensity refers to the power or
loudness of the voice, typically measured in decibels (dB). It correlates perceptually to vocal loudness. It is a crucial parameter in voice assessment because it reflects the efficiency of aerodynamic and phonatory processes involved in voice production. Vocal intensity is determined by subglottal pressure (the pressure below the vocal folds), the resistance offered by the vocal folds to the airflow from the lungs, and the configuration of the vocal tract. For quick reference, Table 6–3 presents average habitual speaking intensity data for adults and children (Kent et al., 2023; Siupsinskiene & Lycke, 2011). Various methods have been proposed for eliciting habitual intensity. Key considerations when measuring intensity are the mouth-to-microphone distance, the level of ambient or background noise, speaking task, and speaking fundamental frequency (SFF). Although no standard exists, a common mouth-to-microphone distance is 12 in. (or 30 cm). One should document the distance and use this consistently when comparing intensity values across sessions. Zraick and colleagues (2004) suggest that clinicians use more than one task to determine habitual loudness. For example, values elicited by having the patient count from 1 to 10, speak spontaneously, and read aloud could be averaged
TABLE 63. Average Speaking Fundamental Frequency (SFF) for Adults and Children
Adult Males Adult Females Children
Mean Range* Mean Range* Mean Range*
Average SFF (Hz)
Note: *Plus or minus two (±2) standard deviations.
112.4
(nonsingers)
130.5
(singers)
89.0–175.0
98.0–175.0
212.4
(nonsingers)
223.6
(singers)
164.5–260.0
181.0–269.0
251.9
(nonsingers)
244.8
(singers)
201.8–302.0
196.0–322.4
180 The Voice and Voice Therapy
before a determination is made about whether therapy to address loudness is necessary. A relatively inexpensive clinical instrument for the measurement of loudness-related parameters is the Level II sound-level meter, purchased from a place such as Sam Ash. Analog and digital versions are available. Most consumer sound-level meters are sensitive from 40 to 130 dB SPL, with slow or fast response for checking peak and average signal levels. The sound-level meter should have the ability to employ different weighting filters, with a C or Linear weighting being the most desirable for voice recordings. Typically, the sound-level meter is held by the clinician at a distance of 30 to 50 cm from the speaker.
Loudness variability: Intensity variability is the range of intensities used in connected
speech. Normal voices have some intensity variability, perceived by the listener as acceptable changes in intonation. In some dysphonic speakers, however, intensity can be either more or less variable than expected or tolerated by the listener. In connected speech, decreased intensity variability may be perceived as monoloudness. Abnormal intensity variability may have either a physiological etiology (such as Parkinson’s disease, vocal fold paralysis, or hearing loss) or may result from learned behavior. Intensity variability is measured in terms of the standard deviation (SD) from the average intensity. This SD reflects the range of intensities around the average intensity,
unemotional sentence is around 10 dB, but it can be higher depending on the speaker’s mood.
Dynamic range: Dynamic range is the physiological range of intensities, from the
softest nonwhisper to the loudest shout, which the patient can produce without undue physical strain. Speakers rarely speak at either end of their dynamic range (approximately 40 to 115 dB) for extended periods. Therefore, the clinician should focus attention on the dynamic range available to the patient around their habitual loudness. Table6–4 presents average dynamic speaking range data for adults and children (Kent et al., 2023; Siupsinskiene & Lycke, 2011). The dynamic range depends on the F0 produced. It tends to be greatest for F0 in the midrange and less for F0 that is much lower or higher (Ferrand, 2007). The fact that F0 and intensity co-vary leads some to propose the use of the voice range profile (VRP) to assess some patients.
TABLE 64. Average Dynamic Speaking Range for Adults and Children
Adult Males Adult Females Children
Mean Range* Mean Range* Mean Range*
Average dynamic speaking range (dBA)
Note: *Plus or minus two (±2) standard deviations.
30.2
(nonsingers)
30.6
(singers)
21.9–38.5 62.1 (nonsingers)
17.4–44.0 61.0
(singers)
55.5–68.6 59.7 (nonsingers)
52.1–69.9 61.5
(singers)
53.0–68.9
56.0–66.9
CHAPTER 6 Evaluation of the Voice 181
Average speaking fundamental frequency (SFF): This measure reflects the average
pitch level used by the speaker for the majority of their vocalizations. It correlates with the auditory perception of habitual pitch. Normative data across the lifespan have been published, and the clinician should use these norms when making a clinical judgment about the suitability of a particular patient’s habitual pitch (Kent et al., 2023; Siupsinskiene & Lycke, 2011). Siupsinskiene and Lycke (2011) compared modal pitch across five different speaking durations (1, 5, 15, 30, and 60 s) and reported significant differences between the 30- and 60-s samples. Zraick, Gentry, and colleagues (2006) compared modal pitch across six different social contexts (speaking during a voice evaluation, speaking in public, speaking to a peer, speaking to a superior, speaking to a subordinate, and speaking to a parent or spouse) and reported that speaking differed depending on who was the patient’s communication partner. Results of these studies (and others, e.g., Sandage et al., 2015) indicate that measures of pitch should be interpreted in light of how it was elicited. A relatively inexpensive clinical instrument for the measurement of pitch-related parameters is the piano or electric keyboard. Isolated vowels or connected speech can be produced by the patient and pitch-matched on the keyboard by the clinician. With a piano or keyboard (or pitch pipe, for that matter), it is possible to estimate SFF because the tones produced by the human voice can be matched to the musical notes of these instruments. For example, in hertz, the typical adult male voice is near C3 (131 Hz), an adult female voice is near A3 (220 Hz), and a child’s voice is between C4 and D4 (262 to 294 Hz). Each octave on a musical instrument is composed of eight whole tones, with each tone represented by an alphabetical letter. Sharps and flats represent semitones. There are 12 semitones in an octave. Each C begins a new octave. Each octave represents a doubling of frequency of vocal fold vibration. Therefore, an increase from C3 to C4 represents a doubling of frequency (131 Hz + 131Hz = 262 Hz). See Table 6–5 for a musical note-to­frequency chart.
Maximum phonational frequency range (MPFR): MPFR refers to the complete range
of pitches that an individual can produce, from the lowest pitch (fundamental frequency) to the highest pitch, using a full voice without falsetto. It is typically measured in hertz (Hz) and can also be expressed in semitones to provide a more intuitive understanding of the vocal range’s musical interval. The MPFR is an important parameter in the assessment of vocal function, as it reflects the flexibility and health of the vocal folds and the efficiency of the vocal tract’s resonating system. Zraick and colleagues (2000) compared two methods for eliciting MPFR (stepping from lowest to highest note versus gliding from lowest to highest note) and reported that stair-step progression through the range resulted in a larger MPFR. Zraick and colleagues (2002) tried to determine whether the lowest or highest pitch should be obtained first and reported that obtaining the lowest pitch followed by the highest pitch resulted in a larger MPFR. Results of these studies and others, such as Ma and Li (2017), indicate that MPFR should be interpreted in light of how it was elicited.
Speaking fundamental frequency variability: This is the range of SFFs used in
connected speech. Normal voices have some frequency variability, perceived by the
182 The Voice and Voice Therapy
TABLE 65. Musical Note-to-Frequency Chart
Note
A
1
B
1
C
2
D
2
E
2
F
2
G
2
A
2
B
2
C
3
D
3
E
3
F
3
G
3
Frequency
(Hz) Note
55 A
62 B
65 C
73 D
82 E
87 F
98 G
110 A
123 B
131 C
147 D
164 E
175 F
196 G
3
3
4
4
4
4
4
4
4
5
5
5
5
5
Frequency
(Hz) Note
220 A
245 B
262 C
294 D
330 E
349 F
392 G
440 A
494 B
523 C
587 D
659 E
698 F
784 G
5
5
6
6
6
6
6
6
6
7
7
7
7
7
Frequency
(Hz)
880
988
1046
1175
1318
1397
1568
1760
1975
2093
2349
2637
7294
3136
listener as acceptable changes in prosody. In some speakers with dysphonia, however, frequency can be either more or less variable than expected or tolerated by the listener. Increased frequency variability may be perceived as a childlike, singsong prosody, while decreased frequency variability may be perceived as monotone. Abnormal frequency variability may have a functional, organic, or neurological basis. For quick reference, Table 6–6 presents average SFF variability data for adults and children (Kent etal., 2023; Siupsinskiene & Lycke, 2011).
Cepstral peak prominence (CPP): CPP is an advanced tool used to check voice quality
by looking at how regular the voice sounds are. In simpler terms, it checks how evenly the vocal folds vibrate. CPP measures the clarity of a specific peak in the voice signal’s cepstrum — a special chart that shows the pattern of the voice’s sound waves (Awan etal., 2010). This peak tells us about the most important sound in the voice. A clear and distinct peak means the voice is steady and smooth, which is a good sign. If the peak is less clear, it might mean there is a voice problem, like hoarseness (Heman-Ackah et al., 2014). CPP is highly recommended for assessing all levels of dysphonia severity, both in sustained vowels and in continuous speech (Maryn et al., 2009). This approach
CHAPTER 6 Evaluation of the Voice 18 3
offers a significant advantage over traditional methods such as jitter and shimmer, which are mainly effective for identifying mild to moderate dysphonia during longer vowel sounds where the speaker tries to maintain a constant pitch and loudness. CPP allows for a broader and more accurate evaluation of voice quality across a range of speaking conditions (Awan et al., 2010). Cepstral measures are available in software programs for clinical use (Murray et al., 2022; Watts et al., 2017). For quick reference, Table 6–7 reports normative CPP data (Buckley et al., 2023; Murton et al., 2020).
Voice range profile: The phonetogram is a graphical representation of an individual’s
vocal SPL plotted against the vocal fundamental frequency (F0) (Sanchez et al., 2014). It is utilized to measure and plot profiles for both speech (speech range profile, SRP) and maximal voice capacity (VRP) (Awan, 1993). The term Voice Range Profile (VRP) was officially proposed by the Voice Committee of the International Association of Logopedics and Phoniatrics in 1992 to describe the span of an individual’s minimum and maximum intensity levels across their entire vocal range. Phonetograms serve as the graphical depictions of the VRP (Figure 6–11), effectively capturing and illustrating the
TABLE 66. Average Speaking Fundamental Frequency Variability (Pitch Sigma) for
Adults and Children
Adult Males Adult Females Children
Mean Range* Mean Range* Mean Range*
Average pitch sigma (ST)
Note: *Plus or minus two (±2) standard deviations.
13.6 11.95–14.17 20.45 19.75–21.15 8.9 4.0–20.7
TA B L E 6 7. Cepstral Peak Prominence (CPP) Norms
Stimulus
ɑ
/, ADSV (CPP) 8.86 dB 8.09 dB
/
ɑ
/, Praat (CPPS) 11.72 dB 11.05 dB
/
/i/, ADSV (CPP) 6.68 dB 4.65 dB
/i/, Praat (CPPS) 12.01 dB 10.37 dB
The Rainbow Passage, ADSV (CPP) 5.40 dB 5.04 dB
The Rainbow Passage, Praat (CPPS) 6.40 dB 6.49 dB
Note: DSV = analysis of dysphonia in speech and voice; CPP/CPPS = cepstral peak prominence/ smoothed cepstral peak prominence.
Normative Value
(Males)
Normative Value
(Females)
184 The Voice and Voice Therapy
FIGURE 613. The Computerized Speech Lab in clinical use. Used with the
permission of Pentax Medical.
dynamic range of vocal capabilities in terms of pitch and loudness, which can be crucial for both clinical assessment and vocal training purposes. To capture a VRP, the speaker is asked to phonate the vowel /i/ or /a/ at select frequencies across their frequency range (modeled by a tone-generator such as a piano or pitch pipe, or presented by computer software), both as softly and loudly as possible. We typically obtain the lower VRP intensity contour before the upper intensity contour to avoid possible laryngeal fatigue, particularly in persons with dysphonia. The following VRP characteristics are usually reported: habitual frequency, total fundamental frequency range, lowest and highest fundamental frequencies, habitual intensity, total intensity range, lowest and highest intensity, and VRP shape and contour. The upper contour of the VRP represents the patient’s maximum phonation threshold (maximum intensity at each frequency), and the lower contour represents their minimum phonation threshold (minimum intensity at each frequency). A normal VRP should have an oblique-oval shape; at the physiological extremes of vocal range, there is a minimal intensity difference between the soft and loud phonations (LeBorgne, 2007). A constricted (compressed) VRP indicates that the patient has difficulty achieving normal frequency and intensity ranges. Obtaining a complete
CHAPTER 6 Evaluation of the Voice 18 5
VRP for some patients can be challenging, leading some to propose various customized VRP protocols for clinical use (Cutchin et al., 2020). Table 6–8 presents VRP data collected on children with normal and dysphonic voices (Andersen et al., 2023; Heylen et al., 1998). Table 6–9 presents VRP data collected on adults with normal and dysphonic voices (Ma et al., 2007). VRP data are also available for male and female professional voice users (Heylen et al., 2002), trained versus untrained singers (Siupsinskiene & Lycke, 2011), and adults with untrained normal voices (Sanchez et al., 2014).
Electroglottogram (EGG): Electroglottography (EGG) is a noninvasive technique for
obtaining an estimate of vocal fold contact patterns during phonation. A gold electrode is placed on each side of the thyroid cartilage at a level corresponding to the position of the vocal folds. A weak high-frequency electrical current is passed between the electrodes. As vocal fold contact area changes, there are changes in the electrical resistance between the electrodes. When the glottis is opening or open, resistance increases; when the glottis is closing or closed, resistance decreases. The resulting Lx waveform, called an electroglottograph or laryngogram, reveals summary information about vocal fold contact over time. The EGG can be used to visualize various types of voice quality. For
TABLE 68. Voice Range Profile Measures for Children With Dysphonic
Voices Versus Children With Normal Voices
Measures Dysphonic† (n = 136) Normal Voice* (n = 94)
Frequency Measures Mean SE Mean SE
Lowest frequency (Hz) 196.3 2.5 192.8 2.5
Highest frequency (Hz) 550.0 11.0 857.0 21.0
Total frequency range (Hz) 354.0 13.0 663.0 22.0
Number of semitones in modal register
Number of semitones in falsetto register
Total number of semitones 19.4 0.5 26.4 0.5
Intensity Measures Mean SE Mean SE
Lowest intensity (in dB) 52.4 0.3 48.2 0.3
Highest intensity (in dB) 95.2 0.5 98.0 0.6
Total intensity range (in dB) 42.7 0.6 49.7 0.6
Note: *Vocally healthy group composed of 53 boys and 41 girls; †Dysphonic group composed of 87 boys and 49 girls; SE = standard errors.
Source: Measures as reported by Heylen and colleagues (1998).
13.8 0.3 18.0 0.3
5.5 0.4 8.4 0.4