CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Peanuttoad /egodistilltextn<1K1 likes8.4k downloads23d agoHugging Face02snehasis19 /opendatalab-experimental-nmr-peaks OpenDataLab Experimental NMR Peaks Dataset Dataset Description This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas. Dataset Summary Total Samples: 533,595 compounds Batches: 333 batch files Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.textother100K<n<1M0 likes2.7k downloads8mo agoHugging Face03Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face04applied-ai-018 /peacock-data-public-datasets-sangrahatext10M<n<100M0 likes774 downloads2y agoHugging Face05peach-lab /CIDER CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment Paper | Code Dataset for the COLM 2026 paper CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment CIDER is a dataset of privacy disclosure decisions collected from real users. It consists of 14,850 annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios. What can you do with… See the full description on the dataset page: https://huggingface.co/datasets/peach-lab/CIDER.imagen<1K0 likes722 downloads21h agoHugging Face06peandrew /conceptnet_en_simpletext1M<n<10M1 likes672 downloads4y agoHugging Face07peakji /peak-anchor-content-35ktabular10K<n<100K0 likes596 downloads2y agoHugging Face08MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes473 downloads10mo agoHugging Face09applied-ai-018 /peacock-data-public-datasetstext1K<n<10K0 likes467 downloads2y agoHugging Face10peakji /peak-search-content-70ktabular10K<n<100K0 likes459 downloads2y agoHugging Face11peakji /peak-intent-50text100K<n<1M0 likes439 downloads2y agoHugging Face12Prosho /pear-data 🍐 PEAR MT Evaluation Data Overview This dataset contains the pairwise Machine Translation evaluation data used to train and evaluate PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation. Each example contains: a source segment; an optional human reference; two candidate translations; the corresponding MT system identifiers; human quality scores for both candidates; contextual metadata such as year, language pair, and domain. The… See the full description on the dataset page: https://huggingface.co/datasets/Prosho/pear-data.tabulartranslation10M<n<100M2 likes360 downloads2mo agoHugging Face13microsoft /PEACE PEACE: Empowering Geologic Map Holistic Understanding with MLLMs [Code] [Paper] [Data] Introduction We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table. Property Description Source USGS(English) CGS(Chinese) Content Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.imagequestion-answering1K<n<10K22 likes353 downloads2y agoHugging Face14peakji /peak-search-300ktabular100K<n<1M0 likes327 downloads2y agoHugging Face15afmck /peanuts-opt-6.7b Peanut Comic Strip Dataset (Snoopy & Co.) This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13. There are 77,457 panels extracted from 17,816 comic strips. The dataset size is approximately 4.4G. Each row in the dataset contains the following fields: image: PIL.Image containing the extracted panel. panel_name: unique identifier for the row. characters: tuple[str, ...] of characters included in the comic strip the panel is part of. themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-opt-6.7b.imagetext-to-image10K<n<100K2 likes314 downloads3y agoHugging Face16PeacefulData /2025_DCASE_AudioQA_Officialgated Audio SFT / Post-Training Data The proposed audio question answering (AQA) dataset with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA) DCASE 2025 Task Description Audio QA Model Baseline Watkins Marine Mammal Sound Database "Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum." 📢 Post-Challenge Research Note While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.text10K<n<100K6 likes275 downloads4mo agoHugging Face17humairmunirawn /and-peaceaudion<1K0 likes268 downloads7mo agoHugging Face18jablonkagroup /nmrexp-cnmr-peaklist-1.5Mtext1M<n<10M0 likes244 downloads29d agoHugging Face19PEARLS-Lab /meow-tea-oolongtext1M<n<10M0 likes232 downloads8mo agoHugging Face20PeacefulData /Neko-v1private for working in progress. ACL 2025 underview. text100K<n<1M0 likes217 downloads1y agoHugging Face21Lihuchen /pearl_benchmark PEARL-Benchmark: A benchmark for evaluating phrase representations Learning High-Quality and General-Purpose Phrase Representations. Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek. Accepted by EACL Findings 2024 Our PEARL Benchmark contains 9 phrase-level datasets of five types of tasks, which cover both the field of data science and natural language processing. Description Paraphrase Classification: PPDB and PPDBfiltered (Wang et al., 2021) Phrase Similarity:… See the full description on the dataset page: https://huggingface.co/datasets/Lihuchen/pearl_benchmark.tabular1M<n<10M1 likes207 downloads3y agoHugging Face22peandrew /conceptnet_en_nomalizedThis is the English part of the ConceptNet and we have removed the useless information. text1M<n<10M2 likes189 downloads4y agoHugging Face23Peanuttoad /StreamGaze_v2 StreamGaze Dataset StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts. Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset. 📁 Dataset Structure streamgaze/ ├── metadata/ │ ├── egtea.csv # EGTEA fixation metadata │ ├── egoexolearn.csv # EgoExoLearn fixation metadata │ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.imagequestion-answering1M<n<10M1 likes180 downloads4mo agoHugging Face24peakji /peak-anchor-40ktabular10K<n<100K0 likes163 downloads2y agoHugging Face25peakji /peak-text-with-context-2mtext1M<n<10M0 likes156 downloads2y agoHugging Face26adityas /PEARLimagen<1K1 likes147 downloads3y agoHugging Face27applied-ai-018 /peacock-data-public-datasets-hubtext100K<n<1M0 likes137 downloads2y agoHugging Face28pearsonkyle /tcg-frame-removal-dataset TCG Frame Removal Dataset 547 paired examples for training instruction-editing models that strip the frame, text, and UI elements from trading-card images and extend the artwork to a seamless full-bleed illustration. This is the training set for the TCG Frame Removal LoRA (FLUX.2-Klein 4B) model (weights). Game Pairs Magic: The Gathering 304 Digimon 154 Pokémon 63 Yu-Gi-Oh! 26 Fields id (string): unique card slug, prefixed by game (mtg-… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/tcg-frame-removal-dataset.imageimage-to-imagen<1K0 likes137 downloads2mo agoHugging Face29deepcopy /PEaCEimage1M<n<10M0 likes136 downloads1y agoHugging Face30Peacockery /tajik-asr-corpus-v3 tajik-asr-corpus-v3 1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled) plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind Peacockery/omni-ctc-300m-tajik (16.9% WER on FLEURS test, 37.6% on held-out conversational speech). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.textautomatic-speech-recognition100K<n<1M3 likes131 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.