CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01CohereLabs /aya_collection_language_split This is a re-upload of the aya_collection, and only differs in the structure of upload. While the original aya_collection is structured by folders split according to dataset name, this dataset is split by language. We recommend you use this version of the dataset if you are only interested in downloading all of the Aya collection for a single or smaller set of languages. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of… See the full description on the dataset page: https://huggingface.co/datasets/CohereLabs/aya_collection_language_split.tabular100M<n<1B122 likes22k downloads1y agoHugging Face02AmelieSchreiber /toricgt-curated-splits ToricGT Curated Graph Reasoning Splits Curated working dataset repository for ToricGT. The upload contains only curated split Parquet files and metadata generated locally. Raw upstream downloads are not uploaded. Each row preserves source dataset, license, split, hashes, and graph JSON fields for audit. Hebrew/Jewish-text records are sourced from Sefaria and UniMorph Hebrew sources. Files train.parquet validation.parquet test.parquet all.parquet if… See the full description on the dataset page: https://huggingface.co/datasets/AmelieSchreiber/toricgt-curated-splits.tabulartext-generation1M<n<10M0 likes1.7k downloads4mo agoHugging Face03nthngdy /wikipedia-22-12-concat-split Dataset Card for "wikipedia-22-12-concat-split" More Information needed tabular10M<n<100M0 likes876 downloads3y agoHugging Face04fdemelo /ipa-childes-split IPA-CHILDES split This dataset is a postprocessed version of the IPA-CHILDES dataset. In particular, the following changes have been implemented: column processed_gloss dropped as it duplicates information of gloss up to punctuation column gloss renamed as sentence, and column ipa_transcription renamed as ipa_g2p_plus (cf. G2P+) column lang added to make IETF language tags accessible for training and inference; language tags normalized by the langcodes package columns ipa_espeak… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ipa-childes-split.tabular10M<n<100M0 likes592 downloads1y agoHugging Face05LocalResearchGroup /split-avelina-python-edutabular1M<n<10M0 likes563 downloads1y agoHugging Face06LocalResearchGroup /split-finemathtabular1M<n<10M0 likes498 downloads1y agoHugging Face07starsofchance /processed_test_splitstabular10K<n<100K0 likes443 downloads1y agoHugging Face08TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptstabular10K<n<100K0 likes424 downloads1y agoHugging Face09finetrainers /OpenVid-60k-split Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M. This is a 60k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered. from datasets import load_dataset, disable_caching, DownloadMode from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-60k-split.tabulartext-to-video10K<n<100K5 likes390 downloads1y agoHugging Face10jwkirchenbauer /fictionalqa_training_splits Training splits view of the FictionalQA dataset The FictionalQA dataset Repository: https://github.com/jwkirchenbauer/fictionalqa Paper: https://arxiv.org/abs/2506.05639 Dataset Description This dataset is a derivative of the main dataset hf.co/datasets/jwkirchenbauer/fictionalqa. Please see that dataset's README for a detailed description of the assets. The dataset splits (configs) provided here are the exact ones materialized and used in the experiments for… See the full description on the dataset page: https://huggingface.co/datasets/jwkirchenbauer/fictionalqa_training_splits.tabulartext-generation100K<n<1M0 likes369 downloads7mo agoHugging Face11Hennara /Recap-DataComp-1B_split_3image100M<n<1B0 likes319 downloads2y agoHugging Face12NomaDamas /split_search_qa preprocessed_SearchQA The SearchQA question-answer pairs originate from J! Archive2, which comprehensively archives all question-answer pairs from the renowned television show Jeopardy! The passages, sourced from Google search web page snippets. We offer passage metadata, encompassing details like 'air_date,' 'category,' 'value,' 'round,' and 'show_number,' enabling you to enhance retrieval performance at your discretion. Should you require further details about SearchQA, please… See the full description on the dataset page: https://huggingface.co/datasets/NomaDamas/split_search_qa.tabular10M<n<100M0 likes316 downloads3y agoHugging Face13Hennara /Recap-DataComp-1B_split_4image100M<n<1B0 likes300 downloads2y agoHugging Face14G4KMU /t2-ragbench-splitstabular10K<n<100K0 likes207 downloads1y agoHugging Face15Hennara /Recap-DataComp-1B_split_5image100M<n<1B0 likes196 downloads2y agoHugging Face16TyroneDragon /highlevel_thinking_with_grounding_annotation_split1000_v2_merged_promptstabular10K<n<100K0 likes175 downloads1y agoHugging Face17nhull /tripadvisor-split-dataset New Version Available A newer version of this dataset with improved annotations and additional examples is available here. tabular10K<n<100K1 likes173 downloads2y agoHugging Face18vibhuiitj /Exercise-Synthetic-split-ncert-chapter-mapped_filtered_difficulty_scoredtabular1M<n<10M0 likes173 downloads5mo agoHugging Face19Hennara /Recap-DataComp-1B_split_7image100M<n<1B0 likes168 downloads2y agoHugging Face20Hennara /Recap-DataComp-1B_split_8image100M<n<1B0 likes152 downloads2y agoHugging Face21Hennara /Recap-DataComp-1B_split_2image100M<n<1B0 likes144 downloads2y agoHugging Face22finetrainers /OpenVid-10k-split Combination of part_id's from bigdata-pw/OpenVid-1M and video data from nkp37/OpenVid-1M. This is a 10k video split of the original dataset for faster iteration during testing. The split was obtained by filtering on aesthetic and motion scores by iteratively increasing their values until there were at most 1000 videos. Only videos containing between 80 and 240 frames were considered. from datasets import load_dataset, disable_caching, DownloadMode from torchcodec.decoders import… See the full description on the dataset page: https://huggingface.co/datasets/finetrainers/OpenVid-10k-split.tabulartext-to-video1K<n<10K2 likes143 downloads1y agoHugging Face23albertge /data_ablation_full59K-modernbert-split-kmeans-dim768-20250218tabular10K<n<100K0 likes140 downloads2y agoHugging Face24windfromthenorth /craft-multiturn-actions-split-nothinktabular1M<n<10M0 likes138 downloads11mo agoHugging Face25BrachioLab /dist-defense-traces-taskname-split-augmented-plus-synth-v15 BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15 Task-name-disjoint train/test splits for dist-defense embedding training. Contents Splits: dist_train, dist_test Built from: output/ctf_packaged_augmented_taskname_split_plus_synth_v15_trainonly Split sizes: dist_train=132231, dist_test=234529 Split params: seed=42, train_ratio=0.9, benign_train_ratio=0.3 Synthetic merge: appended 35891 rows from… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15.tabular100K<n<1M0 likes129 downloads8mo agoHugging Face26Hennara /Recap-DataComp-1B_split_1image100M<n<1B1 likes127 downloads2y agoHugging Face27joekiller /splack-splits-augment-10tabular10M<n<100M0 likes127 downloads10mo agoHugging Face28Hennara /Recap-DataComp-1B_split_6image100M<n<1B0 likes126 downloads2y agoHugging Face29Reza2kn /telephooney-trainability-splits 🗂️ telephooney-trainability-splits English + فارسی · Part of Shenava 1.0 · Project hub · SLT paper submission 🌟 At a glance | معرفی سریع English فارسی 🎯 Purpose Telephooney training split dataset. تقسیم‌بندی‌های آموزش‌پذیری گفتار تلفنی برای تحلیل کیفیت و تصمیم‌گیری دربارهٔ ورود نمونه‌ها به آموزش. 🧩 Role quality calibration and data-selection asset مصنوع کالیبراسیون کیفیت و انتخاب داده 📦 Snapshot 10 files; approximately 299.28 MB 10 فایل؛… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/telephooney-trainability-splits.tabular100K<n<1M1 likes117 downloads2mo agoHugging Face30Artefacts /shape-sorting-so101_split_testThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "observation.state": { "dtype": "float32", "shape": [ 6 ], "names": [ "Rotation", "Pitch", "Elbow", "Wrist_Pitch", "Wrist_Roll", "Jaw" ] }… See the full description on the dataset page: https://huggingface.co/datasets/Artefacts/shape-sorting-so101_split_test.tabularrobotics100K<n<1M0 likes116 downloads7d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.