datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
librispeech-clean-lance
LibriSpeech clean (Lance Format)
A Lance-formatted version of the LibriSpeech ASR clean configuration, sourced from openslr/librispeech_asr. Each row is one utterance with inline FLAC audio bytes, the reference transcript, a sentence-transformers embedding of that transcript, and speaker/chapter metadata — all available directly from the Hub at hf://datasets/lance-format/librispeech-clean-lance/data.
Key features
Inline FLAC bytes in the audio column at 16 kHz mono… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/librispeech-clean-lance.JuzneVesti-SR-Unsloth-Format
Emilia-compatible Serbian speech (JuzneVesti-SR)
This is a format conversion of JuzneVesti-SR v1.0 for Hugging Face audio
training pipelines. It exposes the same columns as
kadirnar/Emilia-DE-B000000 and preserves the original train/dev/test split
(with dev named validation).
Source
Peter Rupnik and Nikola Ljubesic, ASR training dataset for Serbian
JuzneVesti-SR v1.0, Jozef Stefan Institute / CLARIN.SI (2022).
Persistent identifier:… See the full description on the dataset page: https://huggingface.co/datasets/baki83/JuzneVesti-SR-Unsloth-Format.
