Cnam-LMSSC/vibravox-test
Dataset Card for Vibravox-test Important Note This dataset contains a very small proportion (1.2 %) of the original Vibravox Dataset. vibravox-test is a only a dummy dataset for use with test pipelines in the Vibravox project. It is therefore not intended for training or testing models. For full access to the complete dataset and documentation suitable for training and testing various audio and speech-related tasks, please visit the Vibravox Dataset page on… See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/vibravox-test.
Dataset Card for Vibravox-test
Important Note
This dataset contains a very small proportion (1.2 %) of the original Vibravox Dataset.
vibravox-test is a only a dummy dataset for use with test pipelines in the Vibravox project. It is therefore not intended for training or testing models.
For full access to the complete dataset and documentation suitable for training and testing various audio and speech-related tasks, please visit the [Vibravox Dataset](https://huggingface.co/datasets/Cnam-LMSSC/vibravox) page on Hugging Face.
Dataset Details
Structure and Content
The vibravox-test dataset includes a very small fraction of the original [Vibravox Dataset](https://huggingface.co/datasets/Cnam-LMSSC/vibravox) and shares the same structure, comprising:
- speech_clean: Contains clean speech audio samples (0.5% of the original subset). Each split contains 48 rows, with 16 speakers (8 males / 8 females) reading 3 sentences each.
- speech_noisy: Contains noisy speech audio samples. (9.1% of the original subset). Each split contains 48 rows, with 16 speakers (8 males / 8 females) reading 3 sentences each.
- speechless_clean: Contains clean non-speech audio samples. (4.7% of the original subset). Each split contains 3 rows.
- speechless_noisy: Contains noisy non-speech audio samples. (1.5% of the original subset). Each split contains 1 row.
