CoolFace
Datasetpublic

Ugiat/multilingual_librispeech_french_punctuated

Multilingual LibriSpeech French, punctuated and capitalized (train) A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/multilingual_librispeech_french_punctuated.

sourceHugging Facecc-by-4.0updated 3d agoView on Hugging Face
1likes74downloads

Ugiat/multilingual_librispeech_french_punctuated · main · files are served by the source, never re-hosted here