datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
speech_edit_acoustic
SpeechEdit Acoustic Retrieval Dataset
This dataset is an MTEB-formatted Any-to-Any (AT2A) composed audio retrieval adaptation of the acoustic_editing subset of DiscreteSpeech/SpeechEditBench.
Each query combines an original/source speech recording with a natural-language editing instruction, and the corpus contains the corresponding edited target speech recordings.
Schema
queries: id (string), audio (source audio), and text (edit instruction)
corpus: id (string)… See the full description on the dataset page: https://huggingface.co/datasets/deep9539/speech_edit_acoustic.droidnexus-arabic-editorial-speech-scorecard-mini
DroidNexus Arabic Editorial Speech Scorecard Mini
A public DroidNexus Labs scorecard dataset for Arabic speech workflows: representative editorial scenarios, latency targets, overlap pressure, and the metric stack that decides whether a transcript is usable.
Why this exists
This dataset is the first public speech artifact layer for DroidNexus Labs. It publishes representative editorial workloads and evaluation pressure before claiming a full source-audio benchmark.… See the full description on the dataset page: https://huggingface.co/datasets/driodnexus/droidnexus-arabic-editorial-speech-scorecard-mini.
