datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
minimax_music3_qlora_trainerTranscription-Cleanup-Trainer
Text Cleanup Fine-Tuning Dataset
A curated dataset for training speech-to-text cleanup models to achieve optimal transcript refinement.
Dataset Description
This dataset contains paired examples of raw speech-to-text transcriptions and manually-cleaned versions, designed for fine-tuning models to clean up transcripts to a specific quality level ("Goldilocks" cleanup - not too much, not too little).
Dataset Structure
dataset/
├── data/
│ ├── audio/… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Transcription-Cleanup-Trainer.tonic-trainer
Tonic Trainer clips
3094 thirty-second music clips with human-made key annotations, for
ear training. Each clip is the opening 30 seconds of a Free Music Archive
track, re-encoded to mono 128 kbps.
Licensing — read this first
The audio is not ours and is not uniformly licensed. Every clip is
redistributed under the original Creative Commons terms of its own track, which
are recorded per file in attribution.csv. Nothing here is public domain by
default, and this… See the full description on the dataset page: https://huggingface.co/datasets/MrScorcher1/tonic-trainer.asr-interviews-trainer-full
Dataset Card for "asr-interviews-trainer-full"
More Information needed
