Last updated: July 2026
By accessing and using SpeechLab (the "Site"), you agree to the following terms:
SpeechLab is a research demonstration platform providing interactive speech technology demos, including speaker recognition, speech analysis, speech synthesis, and fake speech detection. The Site is intended for educational and research purposes only.
For questions regarding this agreement, please contact the Site administrators through the About page.
Recording requires microphone access. It looks like microphone permission was previously denied for this site.
To enable it:
Meanwhile, you can use the import audio file function.
Fundamental frequency (pitch) tracked over time. Compare contours to observe differences in habitual pitch, intonation patterns, and pitch range. Click voice labels to play the audio.
F1, F2, and F3 are resonant frequencies of the vocal tract that primarily encode vowel quality and provide information about a speaker's vocal tract characteristics.
Long-Term Average Spectrum: mean energy distribution across frequencies (0–6 kHz), computed only over speech-active frames. Mean-normalized to highlight spectral shape differences between voices.
0%
- - - -
Represents the speaker's habitual pitch during the recording.
- - - -
Reflects the degree of pitch variability, providing an indication of intonational variation.
- - - -
Captures the typical span of pitch while reducing the influence of brief extreme values or tracking errors.
- - - -
Cycle-to-cycle variation in pitch period. Lower values indicate a steadier voice.
- - - -
Cycle-to-cycle variation in amplitude. Lower values indicate more stable loudness.
- - - -
Harmonics-to-Noise Ratio. Higher values indicate a clearer, more tonal voice.
- - - -
Cycle-to-cycle variation in pitch period. Lower values indicate a steadier voice.
- - - -
Cycle-to-cycle variation in amplitude. Lower values indicate more stable loudness.
- - - -
Harmonics-to-Noise Ratio. Higher values indicate a clearer, more tonal voice.
- - - -
Hammarberg Index: difference between peak energy in low (0–2 kHz) vs. high (2–5 kHz) frequency bands. Higher values indicate a steeper spectral tilt, typical of more relaxed or breathy voice.
- - - -
Number of voiced segments per second of total audio duration.
- - - -
Percentage of total frames classified as voiced speech.
| Uptime: | - |
|---|---|
| Load (20 s): | - |
| Load (2 min): | - |
| Load (10 min): | - |
| Load (1 h): | - |
| Uptime: | - |
|---|---|
| Load (20 s): | - |
| Load (2 min): | - |
| Load (10 min): | - |
| Load (1 h): | - |
| Uptime: | - |
|---|---|
| Load (20 s): | - |
| Load (2 min): | - |
| Load (10 min): | - |
| Load (1 h): | - |