CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3k downloads3y agoHugging Face02AhmedSaman /tadabur-lora-data-fullaudio10K<n<100K0 likes1.1k downloads27d agoHugging Face03hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes518 downloads15h agoHugging Face04CausalNLP /gpt2small_full_training_datatext1M<n<10M0 likes506 downloads1y agoHugging Face05vargr /yt_full_image_dataset Dataset Card for "yt_full_image_dataset" More Information needed image100K<n<1M1 likes462 downloads3y agoHugging Face06SamuelM0422 /ONS-Capacity-Factor-Dataset-Full Dataset Card for ONS-Capacity-Factor-Dataset-Full Dataset Summary The ONS-Capacity-Factor-Dataset-Full provides hourly data for wind and solar power plants in Brazil. These values are published by the Operador Nacional do Sistema Elétrico (ONS) — the Brazilian National Electric System Operator — which is responsible for coordinating and controlling electricity generation and transmission in the National Interconnected System (SIN). The dataset includes data from 2009 to… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/ONS-Capacity-Factor-Dataset-Full.texttime-series-forecasting10M<n<100M1 likes452 downloads1y agoHugging Face07adrianhenkel /lucidprots_full_data Dataset Card for "lucidprots_full_data" More Information needed 10M<n<100M2 likes425 downloads3y agoHugging Face08spygaurad /BenglaAI_Full_Datasetaudio100K<n<1M0 likes381 downloads3y agoHugging Face09EMBO /soda-vec-data-full_pmc_title_abstract SODA-VEC Clean Dataset This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles. Dataset Overview Total examples: 26,573,900 Training set: 26,473,900 examples (99.6%) Validation set: 50,000 examples (0.2%) Test set: 50,000 examples (0.2%) Quality Filtering Applied This dataset has been processed with the following quality filters: Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.texttext-classification10M<n<100M0 likes363 downloads1y agoHugging Face10aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes330 downloads2mo agoHugging Face11badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes289 downloads11mo agoHugging Face12cryptpesa /anv-data-ke-somali-fullaudio10K<n<100K0 likes284 downloads5mo agoHugging Face13P1ayer-1 /isbndb-full-database Dataset Card for "isbndb-annas" More Information needed text10M<n<100M8 likes201 downloads3y agoHugging Face14acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes165 downloads3y agoHugging Face15pebblebed /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M1 likes154 downloads7mo agoHugging Face16atgarcia /fullDatasetWithSplittext1K<n<10K0 likes153 downloads3y agoHugging Face17quguanni /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M1 likes152 downloads7mo agoHugging Face18snap-stanford /youtube_processed_full_dataset_finaltabular1M<n<10M0 likes149 downloads10mo agoHugging Face19qualiaadmin /full_dataset_grasping-tagged full_dataset_grasping This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics100K<n<1M0 likes144 downloads8mo agoHugging Face20rubricreward /R3-full-datasettext1M<n<10M0 likes142 downloads1y agoHugging Face21didiudom94 /Burn_To_Win_Full-dataset10K<n<100K0 likes128 downloads4mo agoHugging Face22Violet-yo /Chinese-Braille-Dataset-Full-Tone Chinese Braille Sentence Corpus (Full Tone) 📃 [Paper] • 💻 [Code] • 📖 [Passage corpus] • 🎬 [Demo] The sentence-level half of the Braille–Chinese parallel corpus used in "Vision-Braille: A Curriculum Learning Toolkit and Braille–Chinese Corpus for Braille Translation" (EMNLP 2026 Main Conference). Every Braille sequence here retains all tone markers (retention rate r = 100). This is the source corpus: the tone-omission variants used for curriculum training are… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-Full-Tone.texttranslation100K<n<1M2 likes126 downloads25d agoHugging Face23EMBO /soda-vec-data-full_pmc_title_abstract_paired SODA-VEC Paired Dataset for Negative Sampling This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss. Dataset Overview Total examples: 26,573,900 Format: Paired (anchor-positive) for contrastive learning Source: EMBO/soda-vec-data-full_pmc_title_abstract Purpose: Training sentence transformers with negative sampling Data Format Each example contains: anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.textsentence-similarity10M<n<100M1 likes125 downloads1y agoHugging Face24Alexisbo /full_dataset_grasping full_dataset_grasping This dataset was generated using a phospho starter pack. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics100K<n<1M0 likes115 downloads1y agoHugging Face25rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes114 downloads6mo agoHugging Face26closedaxis-12573 /full_pose_retrain_datasetimage100K<n<1M0 likes110 downloads14d agoHugging Face27while0628 /vqasynth_opencv3d_dataset_3600_wotem_v2_fullimage1K<n<10K0 likes107 downloads1y agoHugging Face28davanstrien /dataset-preferences-llm-course-full-dataset Dataset Card for dataset-preferences-llm-course-full-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset.text1K<n<10K1 likes105 downloads2y agoHugging Face29TreeSpecies /species-dataset-full-oakimagen<1K0 likes100 downloads1mo agoHugging Face30rubricreward /R3-full-dataset-no-gluetext1M<n<10M0 likes97 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.