datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
musdb18-hq-flac
MUSDB18-HQ (FLAC Optimized)
Only the audio payload is converted to lossless PCM-16 FLAC. The original columns are preserved: audio, path, and instrument.
Source dataset
This dataset is derived from the original MUSDB18-HQ dataset.
The original dataset card and license are the authoritative references for the source audio and annotations.
Only the audio payload was transcoded to lossless PCM-16 FLAC; paths, instrument labels, and source track structure were… See the full description on the dataset page: https://huggingface.co/datasets/roro128/musdb18-hq-flac.my-voxtral-datasetmalagasy-speech-full_hoby_copyfleurs-flac
FLEURS-FLAC
A losslessly FLAC-compressed version of Google's FLEURS dataset covering 102 languages.
Overview
This repository contains the Google FLEURS dataset repackaged into Parquet shards with PCM24 FLAC-compressed audio binaries.
Key points:
Audio streams are converted to FLAC (PCM24) with sample-level PCM verification against the source.
Sharded into ~500MB Parquet files per split for efficient I/O and streaming.
Covers all 102 languages from the original… See the full description on the dataset page: https://huggingface.co/datasets/roro128/fleurs-flac.Bushiye-Dset_1roryRORA
