CoolFace
20 results

fleurs

google /fleurs FLEURS Fleurs is the speech version of the FLoRes machine translation benchmark. We use 2009 n-way parallel sentences from the FLoRes dev and devtest publicly available sets, in 102 languages. Training sets have around 10 hours of supervision. Speakers of the train sets are different than speakers from the dev/test sets. Multilingual fine-tuning is used and ”unit error rate” (characters, signs) of all languages is averaged. Languages and results are also grouped into seven… See the full description on the dataset page: https://huggingface.co/datasets/google/fleurs.audioautomatic-speech-recognition100K<n<1M468 likes99k downloads4mo agoHugging Facemteb /fleursaudio100K<n<1M0 likes7.1k downloads3mo agoHugging Facegoogle /fleurs-r22 likes4.3k downloads2y agoHugging FaceWueNLP /belebele-fleurs Belebele-Fleurs Belebele-Fleurs is a dataset suitable to evaluate two core tasks: Multilingual Spoken Language Understanding (Listening Comprehension): For each spoken paragraph, the task is to answer a multiple-choice question. The question and four answer choices are provided in text form. Multilingual Long-Form Automatic Speech Recognition (ASR) with Diverse Speakers: By concatenating sentence-level utterances, long-form audio clips (ranging from 30 seconds to 1 minute 30… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/belebele-fleurs.audioaudio-classification10K<n<100K9 likes3.1k downloads2y agoHugging Facebyan /cs-fleurs 🌍 CS-FLEURS: A Massively Multilingual and Code-Switched Speech Dataset 📖 Overview CS-FLEURS is a new dataset for developing and evaluating code-switched speech recognition and translation systems beyond high-resourced languages. 113 unique code-switched language pairs across 52 languages 300 hours of speech data, both read and synthetic 📊 Dataset Statistics CS-FLEURS consists of the following subsets: Read-Test: 14 X-English pairs, read speech… See the full description on the dataset page: https://huggingface.co/datasets/byan/cs-fleurs.audio10K<n<100K19 likes2.7k downloads1y agoHugging FaceWueNLP /sib-fleurs SIB-Fleurs SIB-Fleurs is a dataset suitable to evaluate Multilingual Spoken Language Understanding. For each utterance in Fleurs, the task is to determine the topic the utterance belongs to. The topics are: Science/Technology Travel Politics Sports Health Entertainment Geography Preliminary evaluations can be found at the bottom of the README. The preliminary results in full detail are available in ./results.csv*. Dataset creation This dataset processes and merges… See the full description on the dataset page: https://huggingface.co/datasets/WueNLP/sib-fleurs.audioaudio-classification10K<n<100K15 likes2.4k downloads1y agoHugging Face