CoolFace
Datasetpublic

Ugiat/multilingual_librispeech_french_punctuated

Multilingual LibriSpeech French, punctuated and capitalized (train) A derivative of the French part of Multilingual LibriSpeech (MLS), the corpus of read audiobooks from LibriVox published by Vineel Pratap, Qiantong Xu, Anuroop Sriram, Gabriel Synnaeve and Ronan Collobert (Facebook AI Research). MLS distributes its transcriptions lowercased and without any punctuation. This dataset keeps that upstream text untouched in the text column and adds a second column, text_punctuated… See the full description on the dataset page: https://huggingface.co/datasets/Ugiat/multilingual_librispeech_french_punctuated.

sourceHugging Facecc-by-4.0updated 4d agoView on Hugging Face
1likes74downloads
3 commits on main
be017f14d ago

Add dataset card

RaulQF
b7f593d4d ago

Add files using upload-large-folder tool

RaulQF
fd0daf54d ago

initial commit

RaulQF