datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
StepEval-Audio-Paralinguistic
StepEval-Audio-Paralinguistic Dataset
Paper: Step-Audio 2 Technical ReportCode: https://github.com/stepfun-ai/Step-Audio2Project Page: https://www.stepfun.com/docs/en/step-audio2
Overview
StepEval-Audio-Paralinguistic is a speech-to-speech benchmark designed to evaluate AI models' understanding of paralinguistic information in speech across 11 distinct dimensions. The dataset contains 550 carefully curated and annotated speech samples for assessing capabilities beyond… See the full description on the dataset page: https://huggingface.co/datasets/stepfun-ai/StepEval-Audio-Paralinguistic.ace-step-multilingual
ACE-Step Multilingual Prompt Language Dataset
Dataset accompanying:
"Does Prompt Language Affect AI-Generated Music?
A Multilingual Evaluation of ACE-Step 1.5 Turbo"
Dataset description
250 instrumental tracks generated using ACE-Step 1.5 Turbo.
5 prompt languages
10 musical prompt families
5 matched random seeds
30 seconds per track
48 kHz
identical generation settings across language conditions
Languages
English
Swedish
Spanish
German
French… See the full description on the dataset page: https://huggingface.co/datasets/jacobrrak/ace-step-multilingual.audio-quality-dataset-nfe4-30-step2
Audio Quality Dataset: NFE 4-30 Step 2
Overview
This dataset publishes synthetic speech artifacts and derived spectrograms used for repo-local audio-quality experiments.
At a glance:
2800 synthetic runs
200 short English prompt sentences
14 NFE settings: 4, 6, 8, ..., 30
fixed seed 1024
Each row represents one synthetic run and includes:
prompt text
raw synthetic WAV
processed synthetic WAV
spectrogram PNG
NFE value
procedural weak label
Here, NFE means the number of… See the full description on the dataset page: https://huggingface.co/datasets/TashaSkyUp/audio-quality-dataset-nfe4-30-step2.StepAudio2_Lu_Yin_BaiShiXi_TTSmapalo-stepfun-metadata-real-2AfroRadVoice-FR
AfroRadVoice-FR
Dataset Description
AfroRadVoice-FR is a French speech dataset composed of radiology report recordings, designed to support research in Automatic Speech Recognition (ASR) for African-accented French in medical contexts.
The dataset combines real recordings, synthetic speech, and augmented audio to address data scarcity and improve acoustic diversity in a specialized domain.
Motivation
Current ASR systems show strong performance in… See the full description on the dataset page: https://huggingface.co/datasets/StephaneBah/AfroRadVoice-FR.FormulaEval_datasets
FormulaEval Datasets
This repository provides the official datasets for FormulaEval, a benchmark for evaluating scientific formula vocalization in large speech language models toward accessible learning.
Included Subsets
The dataset repository contains three subsets:
Subset
Domain
Language
Physics700
Physics formulas and equations
Chinese & English (bilingual)
ChemEquation
Chemical equations and formulas
Chinese & English (bilingual)
MixMath
Mathematical… See the full description on the dataset page: https://huggingface.co/datasets/Stephen-Lee/FormulaEval_datasets.swamiji-stepaudio-editx-refsshopwithvoice-yoruba-benchmarkmapalo-stepfun-metadata2dee-stepfun-metadata-real-2trevor-stepfun-metadata2-real-v1step_chemmapalo-stepfun-metadata-real-1voice_Stephen_Frystepik_ml_rustephenfrystephannieDont_be_your_lover_StepAudio_Spoiled_CuteBoy_Spoken_Dataset50-words-stephen-fryDataset of Stephen Fry from his singing on 50 Words For Snow by kate bush.
StepperMotorSoundsWithLabelsCV17_16kHz_es_PROC50-STEPS
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/mmarron14/CV17_16kHz_es_PROC50-STEPS.genshin_StepAudio2_Lu_Yin_answer_Mavuika_audio_samplesStephenSalvatoreOrigin_StepAudio_Spoiled_CuteBoy_Spoken_Datasetfinal-step
