Cnam-LMSSC/non_curated_vibravox
Dataset Card for non-curated VibraVox π This is the non-curated version of the VibraVox dataset. For a full documentation and dataset usage, please refer to https://huggingface.co/datasets/Cnam-LMSSC/vibravox DATASET SUMMARY The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers. This dataset can be used for various audio machine learning tasks : Automatic Speech Recognition (ASR) (Speech-to-Textβ¦ See the full description on the dataset page: https://huggingface.co/datasets/Cnam-LMSSC/non_curated_vibravox.
Dataset Card for non-curated VibraVox
<p align="center"> <img src="https://cdn-uploads.huggingface.co/production/uploads/65302a613ecbe51d6a6ddcec/zhB1fh-c0pjlj-Tr4Vpmr.png" style="object-fit:contain; width:280px; height:280px;" > </p>
π This is the non-curated version of the VibraVox dataset. For a full documentation and dataset usage, please refer to https://huggingface.co/datasets/Cnam-LMSSC/vibravox
DATASET SUMMARY
The VibraVox dataset is a general purpose audio dataset of french speech captured with body-conduction transducers. This dataset can be used for various audio machine learning tasks :
- Automatic Speech Recognition (ASR) (Speech-to-Text , Speech-to-Phoneme)
- Audio Bandwidth Extension (BWE)
- Speaker Verification (SPKV) / identification
- Voice cloning
- etc ...
Citations, links and details
- Homepage: For more information about the project, visit our project page on https://vibravox.cnam.fr
- Github repository: jhauret/vibravox : Source code for ASR, BWE and SPKV tasks using the Vibravox dataset
- Published paper: available (Open Access) on Speech Communication and arXiV
- Point of Contact: Julien Hauret and Γric Bavu
- Curated by: AVA Team of the LMSSC Research Laboratory
- Funded by: Agence Nationale Pour la Recherche / AHEAD Project
- Language: French
- License: Creative Commons Attributions 4.0
If you use the Vibravox dataset (either curated or non-curated versions) for research, cite this paper :
@article{hauret2025vibravox,
title={{Vibravox: A dataset of french speech captured with body-conduction audio sensors}},
author={{Hauret, Julien and Olivier, Malo and Joubaud, Thomas and Langrenne, Christophe and
Poir{\'e}e, Sarah and Zimpfer, V{\'e}ronique and Bavu, {\'E}ric},
journal={Speech Communication},
pages={103238},
year={2025},
publisher={Elsevier}
}and this repository, which is linked to a DOI :
@misc{cnamlmssc2024vibravoxdataset,
author={Hauret, Julien and Olivier, Malo and Langrenne, Christophe and
Poir{\'e}e, Sarah and Bavu, {\'E}ric},
title = { {Vibravox} (Revision 7990b7d) },
year = 2024,
url = { https://huggingface.co/datasets/Cnam-LMSSC/vibravox },
doi = { 10.57967/hf/2727 },
publisher = { Hugging Face }
}