datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KiTS23-sliced-2dlibrispeech_asr_sliced
Librispeech Slices
Description
Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project.
It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz.
A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment.
To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.ship-detection-sliced-bis
Dataset Card for "ship-detection-sliced-bis"
More Information needed
nova-ball-sliced-640Brats2021_slicednova-ball-sliced-320sliced_captionAl-Ahram-raw-sliced-batchsize4000
Dataset Card for "Al-Ahram-raw-sliced-batchsize4000"
More Information needed
260413_r1lite_lerobot_v3_slicedAl-Ahram-raw-sliced
Dataset Card for "Al-Ahram-raw-sliced"
More Information needed
Al-Ahram-raw-sliced-batchsize3000
Dataset Card for "Al-Ahram-raw-sliced-batchsize3000"
More Information needed
sliced_image_caption_v2the_cauldron_ai2d_sliced
How was it built?
dataset = load_dataset("HuggingFaceM4/the_cauldron", "ai2d", split="train", streaming=True)
dataset_iter = iter(dataset)
sliced_dataset = []
for i in range(50):
sliced_dataset.append(next(dataset_iter))
ds = Dataset.from_list(sliced_dataset)
ds.push_to_hub("ariG23498/the_cauldron_ai2d_sliced")
260413_r1lite_lerobot_v3_clean_slicedtask0040_slicedInkuba_xhosa_dev_sliced_v1mri-t1-t2-2D-sliced-64raw_audio_sliced_16khz
Armenian 7500-Hour Raw Audio Sliced (Gated Public)
This dataset contains ~7500 hours of unlabelled Armenian speech processed via Voice Activity Detection (VAD) and downsampled to 16 kHz (Mono, PCM_16).
Packaged into snappy-compressed Parquet shards of ~500 MB for direct compatibility with Hugging Face datasets and high-speed SSL pre-training (HuBERT, Wav2Vec2).
Access is Gated: Manual approval required.
Inkuba_xhosa_train_sliced_v1Inkuba_isizulu_train_sliced_v1Inkuba_isizulu_dev_sliced_v1Inkuba_english_train_sliced_v1ship-detection-sliced-yolo-format-train-onlydcug_md_sliced_829Inkuba_english_dev_sliced_v1CNPM_slicedMixed_dataset_speech_sliced_clean
