Speech recordings by Cancan Li and colleagues.
Released under CC BY-NC-SA 4.0 for noncommercial research.
Background noise from the corpus by David Snyder, Guoguo Chen, and Daniel Povey.
The corpus is released under CC BY 4.0. The FreeSound noise recordings used here are marked public domain; individual SoundBible credits appear below.
Environmental noise recordings by Joachim Thiemann, Nobutaka Ito, and Emmanuel Vincent.
Released under CC BY-SA 3.0, as stated in the dataset description.
Audio and video processing
Noisy inputs contain additive noise at the indicated SNR. Converted speech is produced by the named systems. Lip videos show grayscale mouth crops from the original whispered utterances and are not retimed to synthesized speech.
Original SoundBible recordings
The following recordings are included through MUSAN. Attribution licenses link to CC BY 3.0 where applicable; recordings marked public domain retain that status.