datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vctk_dataset_no_unknownso100_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 22590,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/UN-kk/so100_test.ruv_tv_unknown_speakersDataset copied from http://hdl.handle.net/20.500.12537/191 by Reykjavik University.
Information can be found at that link.
RUV TV unknown speakers
About the RUV TV unknown speakers corpus
The RUV TV unknown speakers corpus is 281 hours of TV data from six RÚV TV
shows. The data continas 221,759 utterrances from various unlabelled speakers.
The text is normalized. The data is aligned and segmented, ready for ASR
training. Audio conditions vary between recordings. This data set is… See the full description on the dataset page: https://huggingface.co/datasets/tiro-is/ruv_tv_unknown_speakers.wikiart-artist-unknown-artistinto-the-unknownvctk_dataset_no_unknown_splittedTDevilsUnKEBenchcombined-unknown-pneumonia-and-tuberculosisicrm-hitek-full-db-mixed
ICMR + HITEK Full DB (Mixed) — Prebuilt Indexes + One-Click Setup
Prebuilt sorted indexes for the Kzr0xx/Icmr-and-hitek dataset (2.5B rows, 11 columns, ~104 GB raw parquet).
Building these indexes took ~17 hours of compute. This repo saves you that work: download + run = API live in ~1-2 hours (download speed dependent).
Contents
The indexes are stored as sorted parts (each < 50 GB, split at row-group boundaries, order preserved) because HuggingFace's classic HTTP… See the full description on the dataset page: https://huggingface.co/datasets/unknownlll2829/icrm-hitek-full-db-mixed.hotpot_qa_unknownALCUNA_meta_affirmative_known_unknownrecord-firstThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 21487,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/UN-kk/record-first.Phi3_intent_v37_2_wo_unknownGV_Train_100h_UnknownGenderPhi3_intent_v45_1_w_unknown_upper_lowerPhi3_intent_v57_2_w_unknown_upper_lowerPhi3_intent_v45_3_w_unknownPhi3_intent_v46_2_w_unknown_upper_lowerPhi3_intent_v68_2_w_unknown_upper_lowerPhi3_intent_v50_2_w_unknown_upper_lowerafrica-worldbank-repeaters-in-grade-unknown-of-lower-secondary-general-education-male-number-uis
Repeaters in grade unknown of lower secondary general education, male (number) | Africa (World Bank — Education Statistics) | Africa (World Bank)
Size category: n<1K - Formats: parquet - Sector: education - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-worldbank-repeaters-in-grade-unknown-of-lower-secondary-general-education-male-number-uis.GV_Train_100h_UnknownGender_qualityKnown-to-Unknown-SQuADv2Phi3_intent_v46_1_w_unknownPhi3_intent_v47_3_w_unknownALCUNA_meta_affirmative_3_mix_position_known_unknown_trainbuzz_sources_125_unknownALCUNA_meta_affirmative_known_unknown_boringALCUNA_meta_affirmative_known_unknown_boring_for_fix_middle_train
