CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01acmc /beamit-full-texts-dataset Dataset Card for "beamit-full-texts-dataset" More Information needed text10K<n<100K0 likes3.5k downloads3y agoHugging Face02AhmedSaman /tadabur-lora-data-fullaudio10K<n<100K0 likes1.1k downloads28d agoHugging Face03CausalNLP /gpt2small_full_training_datatext1M<n<10M0 likes560 downloads1y agoHugging Face04hulk10 /data_gouv_datasets_catalog-full-documents 🇫🇷 Catalogue des jeux de données de data.gouv.fr – Version structurée Ce dataset constitue une version structurée et exhaustive du catalogue des jeux de données publiés sur data.gouv.fr, la plateforme nationale française de l’open data. Il recense l’ensemble des jeux de données référencés sur la plateforme et fournit leurs métadonnées complètes : titre et description, organisation productrice, licence, couverture spatiale et temporelle, fréquence de mise à jour, formats… See the full description on the dataset page: https://huggingface.co/datasets/hulk10/data_gouv_datasets_catalog-full-documents.tabular100K<n<1M0 likes511 downloads16h agoHugging Face05vargr /yt_full_image_dataset Dataset Card for "yt_full_image_dataset" More Information needed image100K<n<1M1 likes510 downloads3y agoHugging Face06SamuelM0422 /ONS-Capacity-Factor-Dataset-Full Dataset Card for ONS-Capacity-Factor-Dataset-Full Dataset Summary The ONS-Capacity-Factor-Dataset-Full provides hourly data for wind and solar power plants in Brazil. These values are published by the Operador Nacional do Sistema Elétrico (ONS) — the Brazilian National Electric System Operator — which is responsible for coordinating and controlling electricity generation and transmission in the National Interconnected System (SIN). The dataset includes data from 2009 to… See the full description on the dataset page: https://huggingface.co/datasets/SamuelM0422/ONS-Capacity-Factor-Dataset-Full.texttime-series-forecasting10M<n<100M1 likes452 downloads1y agoHugging Face07spygaurad /BenglaAI_Full_Datasetaudio100K<n<1M0 likes436 downloads3y agoHugging Face08EMBO /soda-vec-data-full_pmc_title_abstract SODA-VEC Clean Dataset This is a cleaned and filtered version of the SODA-VEC dataset, containing high-quality biomedical title-abstract pairs from PubMed Central (PMC) articles. Dataset Overview Total examples: 26,573,900 Training set: 26,473,900 examples (99.6%) Validation set: 50,000 examples (0.2%) Test set: 50,000 examples (0.2%) Quality Filtering Applied This dataset has been processed with the following quality filters: Abstract Length… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract.texttext-classification10M<n<100M0 likes419 downloads1y agoHugging Face09cryptpesa /anv-data-ke-somali-fullaudio10K<n<100K0 likes332 downloads5mo agoHugging Face10badrex /anv-data-ke-somali-fullaudio100K<n<1M1 likes326 downloads11mo agoHugging Face11aiacademy-kg /house_kg_full_dataset house.kg — Kyrgyzstan Real Estate (multimodal) A complete snapshot of house.kg, the largest real-estate board in Kyrgyzstan: every sale and rental listing, with coordinates, prices, seller identities, agency ratings, reviews — and 227,294 photographs. Field names are English; values are kept in the original language (Russian/Kyrgyz), exactly as the site renders them. 💻 Scraper source code on GitHub → The complete, open scraper that produced this dataset —… See the full description on the dataset page: https://huggingface.co/datasets/aiacademy-kg/house_kg_full_dataset.imagetabular-regression100K<n<1M0 likes317 downloads2mo agoHugging Face12P1ayer-1 /isbndb-full-database Dataset Card for "isbndb-annas" More Information needed text10M<n<100M8 likes217 downloads3y agoHugging Face13acmc /beamit-annotated-full-texts-dataset Dataset Card for "beamit-annotated-full-texts-dataset" More Information needed tabular10K<n<100K0 likes200 downloads3y agoHugging Face14rubricreward /R3-full-datasettext1M<n<10M0 likes175 downloads1y agoHugging Face15atgarcia /fullDatasetWithSplittext1K<n<10K0 likes170 downloads3y agoHugging Face16pebblebed /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/pebblebed/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M1 likes169 downloads7mo agoHugging Face17snap-stanford /youtube_processed_full_dataset_finaltabular1M<n<10M0 likes155 downloads10mo agoHugging Face18quguanni /kernel-vuln-dataset-full Linux Kernel Vulnerability-Introducing Commits Dataset Dataset Description A labeled dataset of 1,426,202 Linux kernel git commits with full metadata, diffs, and binary labels indicating whether each commit introduced a vulnerability that was later fixed. Intended use: Training and evaluating models for vulnerability-introducing commit detection — predicting whether a given code change will later require a security or bug fix. How the Data Was Collected… See the full description on the dataset page: https://huggingface.co/datasets/quguanni/kernel-vuln-dataset-full.tabulartext-classification1M<n<10M1 likes154 downloads7mo agoHugging Face19EMBO /soda-vec-data-full_pmc_title_abstract_paired SODA-VEC Paired Dataset for Negative Sampling This is a paired version of the SODA-VEC dataset, specifically formatted for negative sampling training with MultipleNegativesRankingLoss. Dataset Overview Total examples: 26,573,900 Format: Paired (anchor-positive) for contrastive learning Source: EMBO/soda-vec-data-full_pmc_title_abstract Purpose: Training sentence transformers with negative sampling Data Format Each example contains: anchor (string): The title… See the full description on the dataset page: https://huggingface.co/datasets/EMBO/soda-vec-data-full_pmc_title_abstract_paired.textsentence-similarity10M<n<100M1 likes129 downloads1y agoHugging Face20Violet-yo /Chinese-Braille-Dataset-Full-Tone Chinese Braille Sentence Corpus (Full Tone) 📃 [Paper] • 💻 [Code] • 📖 [Passage corpus] • 🎬 [Demo] The sentence-level half of the Braille–Chinese parallel corpus used in "Vision-Braille: A Curriculum Learning Toolkit and Braille–Chinese Corpus for Braille Translation" (EMNLP 2026 Main Conference). Every Braille sequence here retains all tone markers (retention rate r = 100). This is the source corpus: the tone-omission variants used for curriculum training are… See the full description on the dataset page: https://huggingface.co/datasets/Violet-yo/Chinese-Braille-Dataset-Full-Tone.texttranslation100K<n<1M2 likes127 downloads26d agoHugging Face21rubricreward /R3-full-dataset-no-gluetext1M<n<10M0 likes127 downloads1y agoHugging Face22while0628 /vqasynth_opencv3d_dataset_3600_wotem_v2_fullimage1K<n<10K0 likes114 downloads1y agoHugging Face23closedaxis-12573 /full_pose_retrain_datasetimage100K<n<1M0 likes111 downloads15d agoHugging Face24davanstrien /dataset-preferences-llm-course-full-dataset Dataset Card for dataset-preferences-llm-course-full-dataset This dataset has been created with distilabel. Dataset Summary This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI: distilabel pipeline run --config "https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset/raw/main/pipeline.yaml" or explore the configuration: distilabel pipeline… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/dataset-preferences-llm-course-full-dataset.text1K<n<10K1 likes105 downloads2y agoHugging Face25TreeSpecies /species-dataset-full-oakimagen<1K0 likes100 downloads1mo agoHugging Face26glenn2 /legemma_t2t_data_544k_fullaudio100K<n<1M0 likes88 downloads2y agoHugging Face27rescommons /Full-Ecom-Chatbot-Dataset E-commerce Chatbot Training Data A curated, multi-source dataset for training and evaluating e-commerce conversational AI systems. It covers a broad range of customer intents — from product discovery and order management to returns, tool-augmented responses, and RAG-grounded Q&A — across 16+ product domains. Dataset Summary Split Records Train 35,213 Test 8,818 Total 44,031 The train/test split uses prompt-group-level stratified sampling on source ×… See the full description on the dataset page: https://huggingface.co/datasets/rescommons/Full-Ecom-Chatbot-Dataset.tabularquestion-answering10K<n<100K0 likes88 downloads6mo agoHugging Face28taufiqsyed /salami_data_cleaned_fullaudio10K<n<100K0 likes87 downloads2y agoHugging Face29datapointai /text-2-image-dpo-human-preferences-fullgated Text-2-Image DPO Human Preferences (Full) The complete human preference dataset for text-to-image generation. 416,360 pairwise judgments from ~20,000 annotators comparing AI-generated images across two evaluation dimensions: prompt alignment and overall preference. This is the full, unfiltered version with uniform vote weights. For quality-filtered subsets with calibrated annotator weighting, see: datapointai/text-2-image-dpo-human-preferences (5,000 pairs, trust-weighted)… See the full description on the dataset page: https://huggingface.co/datasets/datapointai/text-2-image-dpo-human-preferences-full.imageimage-classification10K<n<100K1 likes86 downloads6mo agoHugging Face30syvb /nanonla-qwen3-8b-L24-data-full Qwen3-8B NLA — FULL parquets (activation_vector regenerated) The slim NLA splits with the activation_vector column recomputed (raw layer-24 residual at the final token of detokenized_text_truncated). Three configs: av_sft / ar_sft (warm-start SFT) and rl (RL + held-out eval). Each has a different prompt schema, hence separate configs. tabular100K<n<1M1 likes85 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.