CoolFace
13 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ARTPARK-IISc /VaanigatedVAANI is an India-representative multi-modal multi-lingual dataset. The current version (phase 1- 80 districts, phase 2- 85 districts) contains ~31278 hours of spontaenous,image-prompted speech by 156K speakers across 165 districts, talking about 288K images covering 105 languages. From this audio data, 2,122 hours of transcribed data(text) is available, spanning almost evenly across the 165 districts. Project Vaani, by IISc, Bangalore and ARTPARK, is capturing the true diversity of India’s… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani.audioautomatic-speech-recognition1M<n<10M155 likes19k downloads5d agoHugging Face02ArtificialAnalysis /Earnings22-Cleaned-AA Earnings22-Cleaned-AA Quick links: AA Speech-to-Text Leaderboard | AA-WER v2.0 article Earnings22-Cleaned-AA is a cleaned subset of the English Earnings-22 test data from esb/datasets, a corpus of corporate earnings calls from global companies with speakers of many different nationalities and accents. This cleaned subset is the Earnings-22 portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA.audioautomatic-speech-recognitionn<1K6 likes6.4k downloads7mo agoHugging Face03ARTPARK-IISc /Vaani-transcription-partgatedThis dataset is part of the Vaani dataset and consists of only transcribed speech data. It has a total duration of 2041.54 hours, covering 59 languages. This table represents the audio and transcription duration data for various languages. Language Angami Angika Ao Assamese Awadhi Bajjika Bearybashe Bengali Bhili Bhojpuri Bundeli Chakhesang Chakma Chhattisgarhi English Garhwali Garo Gondi Gujarati Halbi Haryanvi Hindi IduMishmi Kannada Kashmiri Karbi Khariboli Khortha Kokborok Konkani… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-transcription-part.audioautomatic-speech-recognition1M<n<10M20 likes1.6k downloads6mo agoHugging Face04ArtificialAnalysis /VoxPopuli-Cleaned-AA VoxPopuli-Cleaned-AA Quick links: AA Speech to Text Leaderboard | AA-WER v2.0 article VoxPopuli-Cleaned-AA is a cleaned subset of the English VoxPopuli test data from esb/datasets, a speech dataset derived from European Parliament recordings. This cleaned subset is the VoxPopuli portion included in AA-WER v2. We manually reviewed and corrected errors in the original ground-truth transcriptions to ensure fairer evaluation of Speech to Text (STT) models. This dataset is part of AA-WER… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/VoxPopuli-Cleaned-AA.audioautomatic-speech-recognitionn<1K7 likes1.1k downloads7mo agoHugging Face05ArtificialAnalysis /Earnings22-Cleaned-AA-chunked Earnings22-Cleaned-AA-chunked Quick links: AA Streaming Speech to Text Leaderboard | Speech to Text methodology Earnings22-Cleaned-AA-chunked is a chunked version of Earnings22-Cleaned-AA, the cleaned Earnings-22 subset used by Artificial Analysis for streaming Speech to Text evaluation. The original Earnings-22 data comes from esb/datasets, a corpus of corporate earnings calls. Artificial Analysis manually reviewed and corrected the reference transcripts in the cleaned subset… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/Earnings22-Cleaned-AA-chunked.audioautomatic-speech-recognitionn<1K1 likes550 downloads3mo agoHugging Face06coml /vox-communis-artifacts VoxCommunis Artifacts The data artifacts used to train and evaluate the MauBERT models, introduced by the MauBERT paper: manifests, frame-level phone alignments, per-language phone inventories, and the language table. Companion model repositories: coml/maubert-feat — articulatory feature prediction. coml/maubert-phone — IPA phone prediction. No audio is included The recordings come from Common Voice 16.1, which you must download yourself; the phone annotations… See the full description on the dataset page: https://huggingface.co/datasets/coml/vox-communis-artifacts.textautomatic-speech-recognition1M<n<10M0 likes273 downloads2mo agoHugging Face07baryonlabs /open-ko-s2s-eval-artifacts Open Ko-S2S 평가 산출물 (감사용) ⚠️ KsponSpeech 참조 전사는 해시로 대체돼 있습니다 KsponSpeech 는 AI Hub 배포 데이터로 재배포 제한이 있을 수 있어, kspon 런의 ref 컬럼을 ref_sha256 으로 대체했습니다(전사 원문 미포함). 모델 출력(hyp)과 채점 결과(cer_err/cer_len/cer)는 우리 산출물이라 그대로 공개합니다. Zeroth 런은 원본이 CC BY 4.0(OpenSLR #40)이라 ref 원문을 그대로 담고 있습니다. 라이선스 보유자의 검증 절차 AI Hub 에서 KsponSpeech 를 정당하게 받은 분은 다음으로 우리 수치를 검증할 수 있습니다. 리더보드 저장소의 eval/datasets_ko.py 에서 clean_kspon() 을 가져옵니다. 자기 사본의 원 전사에 clean_kspon() 을 적용합니다. 결과가 목록이면… See the full description on the dataset page: https://huggingface.co/datasets/baryonlabs/open-ko-s2s-eval-artifacts.tabularautomatic-speech-recognitionn<1K0 likes190 downloads1mo agoHugging Face08ARTPARK-IISc /Vaani-Benchmark-V1.0gated Vaani-Benchmark-V1.0 A curated ASR evaluation set drawn from the Vaani project. This benchmark contains 5,050 audio segments from 1,103 speakers across 104 Indian districts, each with three independent human transcriptions. Evaluation Toolkit A standalone toolkit implementing this benchmark's scoring methodology, plus Latin-script normalization for code-switched predictions and one-command publishing of results to a model's HF card, is available at… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Benchmark-V1.0.audioautomatic-speech-recognition1K<n<10K7 likes181 downloads1mo agoHugging Face09artem-tish /veche-bench О датасете Открытый русскоязычный датасет для оценки систем автоматического протоколирования совещаний: транскрибации (ASR), диаризации и суммаризации. Записи собраны по заранее подготовленным сценариям, озвученным добровольцами, и сопровождаются эталонной разметкой на всех трёх этапах. Исходный код сбора данных, постобработки и бенчмарка: github.com/veche-bench. Состав Совещаний 9 Реплик 1 290 Говорящих (ролей) 54 (от 2 до 10 на встречу)… See the full description on the dataset page: https://huggingface.co/datasets/artem-tish/veche-bench.audioautomatic-speech-recognition1 likes61 downloads3mo agoHugging Face10arthoho66 /medicineaudioautomatic-speech-recognitionn<1K0 likes28 downloads3y agoHugging Face11DALiH-ANR /artsakh_hy Artsakh Dataset (Audio–Transcription Alignment) An audio–transcription alignment dataset for the Artsakh varieties (Stepanakert, Getashen, Hadrut) of Armenian. Overview   This dataset contains Armenian dialect speech recordings (Artsakh varieties: Stepanakert, Getashen, Hadrut) paired with aligned transcriptions. It is split into three subsets:   Train: 5,122 examples  Validation (dev): 140 examples  Test: 175 examples Content   Each example includes:… See the full description on the dataset page: https://huggingface.co/datasets/DALiH-ANR/artsakh_hy.audioautomatic-speech-recognition1K<n<10K2 likes22 downloads7mo agoHugging Face12artmelancholy /golos_mfa_punctuation_long Golos MFA Punctuation (Long) Long-form Russian speech derived from govnejri/golos_mfa_punctuation. Purpose Most public Russian STT corpora ship as short clips (a few seconds each). For benchmarking long-form transcription, VAD, punctuation, and streaming behavior, you want minutes-long audio with reliable word-level alignments. This dataset builds those long clips by splicing groups of consecutive short clips together, inserting randomized silences between them, and… See the full description on the dataset page: https://huggingface.co/datasets/artmelancholy/golos_mfa_punctuation_long.audioautomatic-speech-recognition1K<n<10K1 likes18 downloads4mo agoHugging Face13ARTPARK-IISc /Vaani-Atypical-Speech-CorpusgatedProject Euphonia is a public initiative led by Google that aims to improve Automatic Speech Recognition (ASR) for individuals with atypical speech. To date, most of Project Euphonia’s work has focused on English, resulting in outcomes such as the Android application Project Relate, which generates personalized speech recognition models in English. In recent years, the project has expanded its data collection efforts to additional languages, including French, Spanish, Japanese, and Hindi. The… See the full description on the dataset page: https://huggingface.co/datasets/ARTPARK-IISc/Vaani-Atypical-Speech-Corpus.audioautomatic-speech-recognition1K<n<10K2 likes17 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.