CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01javi22 /high_quality_spanish_speechaudio10K<n<100K1 likes41 downloads8mo agoHugging Face02noystl /speech-quality-datasetThis dataset is a curated subset of 631 speeches selected from the ibm-research/debate_speeches corpus. In our work, Debatable Intelligence: Benchmarking LLM Judges via Debate Speech Evaluation, we use this subset to benchmark LLM judges on the task of debate speech evaluation. Data fields id: The unique identifier for the speech. topic_id: The unique identifier for the topic. topic: The topic of the debate speech (e.g., "Community service should be mandatory"). source: The speech source… See the full description on the dataset page: https://huggingface.co/datasets/noystl/speech-quality-dataset.n<1K0 likes17 downloads1y agoHugging Face03PeacefulData /speech-quality-descriptive-captiongated Audio Large Language Models Can Be Descriptive Speech Quality Evaluators This repository contains code for generating descriptive captions for speech quality evaluation based on the paper "Audio Large Language Models Can Be Descriptive Speech Quality Evaluators" (ICLR 2025). The framework of ALLD and training examples. “Meta info.” is the multi-dimensional ratings annotated by human listeners for the pairwise speech sample. ALLD aims to align the audio LLM response ya to… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/speech-quality-descriptive-caption.text10K<n<100K7 likes16 downloads1y agoHugging Face04BophaAI /low_quality_khmer_speechaudion<1K0 likes13 downloads5mo agoHugging Face05DataoceanAI /Ten_Thousand_People_Dialect_with_High_Quality_Labeling_Speech_Corpus SPECIFICATION: This dataset covers 29,954 dialect speakers from 26 provinces in China, ranging in age from 12 to 75, with a total recording time of 34,073 hours and an average recording duration of nearly 60 minutes, maintaining a balanced gender ratio. The topics covered are very extensive, including news, text messages, vehicle control, music, general, maps, daily colloquial speech, family, health, travel, work, socializing, celebrities, weather, and other common life topics. For… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/Ten_Thousand_People_Dialect_with_High_Quality_Labeling_Speech_Corpus.0 likes5 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.