CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01snehasis19 /opendatalab-experimental-nmr-peaks OpenDataLab Experimental NMR Peaks Dataset Dataset Description This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas. Dataset Summary Total Samples: 533,595 compounds Batches: 333 batch files Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.textother100K<n<1M0 likes2.7k downloads8mo agoHugging Face02Peacockery /common-voice-scripted-speech-26 Common Voice Scripted Speech A row-normalized multilingual ASR dataset built from Mozilla Data Collective Common Voice Scripted Speech. Each upstream archive is converted to appendable parquet shards under data/<upstream_split>/, one shard per source archive and split, with audio bytes embedded in an audio struct column. Status Manifest languages: 60 Languages uploaded: 18 Columns audio (bytes, path) sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.tabularautomatic-speech-recognition100K<n<1M0 likes2k downloads3mo agoHugging Face03peandrew /conceptnet_en_simpletext1M<n<10M1 likes672 downloads4y agoHugging Face04peakji /peak-anchor-content-35ktabular10K<n<100K0 likes596 downloads2y agoHugging Face05MohamedRashad /Arabic-VLM-Full-Pearl 💎 The Arabic VLM Dataset (Full Pearl Edition) This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper. Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.imagequestion-answering100K<n<1M10 likes473 downloads10mo agoHugging Face06peakji /peak-search-content-70ktabular10K<n<100K0 likes459 downloads2y agoHugging Face07peakji /peak-intent-50text100K<n<1M0 likes439 downloads2y agoHugging Face08Prosho /pear-data 🍐 PEAR MT Evaluation Data Overview This dataset contains the pairwise Machine Translation evaluation data used to train and evaluate PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation. Each example contains: a source segment; an optional human reference; two candidate translations; the corresponding MT system identifiers; human quality scores for both candidates; contextual metadata such as year, language pair, and domain. The… See the full description on the dataset page: https://huggingface.co/datasets/Prosho/pear-data.tabulartranslation10M<n<100M2 likes360 downloads2mo agoHugging Face09peakji /peak-search-300ktabular100K<n<1M0 likes327 downloads2y agoHugging Face10afmck /peanuts-opt-6.7b Peanut Comic Strip Dataset (Snoopy & Co.) This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13. There are 77,457 panels extracted from 17,816 comic strips. The dataset size is approximately 4.4G. Each row in the dataset contains the following fields: image: PIL.Image containing the extracted panel. panel_name: unique identifier for the row. characters: tuple[str, ...] of characters included in the comic strip the panel is part of. themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-opt-6.7b.imagetext-to-image10K<n<100K2 likes314 downloads3y agoHugging Face11humairmunirawn /and-peaceaudion<1K0 likes268 downloads7mo agoHugging Face12jablonkagroup /nmrexp-cnmr-peaklist-1.5Mtext1M<n<10M0 likes244 downloads29d agoHugging Face13PeacefulData /Neko-v1private for working in progress. ACL 2025 underview. text100K<n<1M0 likes217 downloads1y agoHugging Face14peandrew /conceptnet_en_nomalizedThis is the English part of the ConceptNet and we have removed the useless information. text1M<n<10M2 likes189 downloads4y agoHugging Face15peakji /peak-anchor-40ktabular10K<n<100K0 likes163 downloads2y agoHugging Face16peakji /peak-text-with-context-2mtext1M<n<10M0 likes156 downloads2y agoHugging Face17applied-ai-018 /peacock-data-public-datasets-hubtext100K<n<1M0 likes137 downloads2y agoHugging Face18pearsonkyle /tcg-frame-removal-dataset TCG Frame Removal Dataset 547 paired examples for training instruction-editing models that strip the frame, text, and UI elements from trading-card images and extend the artwork to a seamless full-bleed illustration. This is the training set for the TCG Frame Removal LoRA (FLUX.2-Klein 4B) model (weights). Game Pairs Magic: The Gathering 304 Digimon 154 Pokémon 63 Yu-Gi-Oh! 26 Fields id (string): unique card slug, prefixed by game (mtg-… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/tcg-frame-removal-dataset.imageimage-to-imagen<1K0 likes137 downloads2mo agoHugging Face19deepcopy /PEaCEimage1M<n<10M0 likes136 downloads1y agoHugging Face20Peacockery /tajik-asr-corpus-v3 tajik-asr-corpus-v3 1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled) plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind Peacockery/omni-ctc-300m-tajik (16.9% WER on FLEURS test, 37.6% on held-out conversational speech). Layout Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/. Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list), and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.textautomatic-speech-recognition100K<n<1M3 likes131 downloads4mo agoHugging Face21PeacefulData /HyPoradise-pilot Dataset Name: Pilot dataset for Multi-domain ASR corrections Description This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains. It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0 Structure Data Split The dataset is divided into training and test splits: Training Data: 281,082 entries Approximately… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HyPoradise-pilot.text100K<n<1M2 likes124 downloads2y agoHugging Face22wjbmattingly /peabody-testimage1K<n<10K0 likes121 downloads11mo agoHugging Face23deepcopy /PEaCE-512pximage1M<n<10M0 likes119 downloads1y agoHugging Face24willxxy /ecg-comprehension-r-peak-count-789480-10000-250-2500text100K<n<1M0 likes116 downloads7mo agoHugging Face25afmck /peanuts-flan-t5-xl Peanut Comic Strip Dataset (Snoopy & Co.) This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13. There are 77,456 panels extracted from 17,816 comic strips. The dataset size is approximately 4.4G. Each row in the dataset contains the following fields: image: PIL.Image containing the extracted panel. panel_name: unique identifier for the row. characters: tuple[str, ...] of characters included in the comic strip the panel is part of. themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-flan-t5-xl.imagetext-to-image10K<n<100K6 likes113 downloads3y agoHugging Face26salahkhenfer /pear-pests-dataset 🍐 Pear Leaf Pest & Disease Detection Dataset This dataset supports object detection for pear leaf health monitoring: locating and classifying an insect pest and fungal/bacterial disease lesions directly on pear leaf images. It targets the timely detection and localization of foliar pear pests and diseases central to precision agriculture, where manual agronomist inspection is labor-intensive, subjective, and hard to scale. The dataset consists of 2,210 annotated images (1,542… See the full description on the dataset page: https://huggingface.co/datasets/salahkhenfer/pear-pests-dataset.imageobject-detection1K<n<10K0 likes108 downloads1mo agoHugging Face27PeacefulData /SINE SINE Dataset Overview The Speech INfilling Edit (SINE) dataset is a comprehensive collection for speech deepfake detection and audio authenticity verification. This dataset contains ~87GB of audio data distributed across 32 splits, featuring both authentic and synthetically manipulated speech samples. Dataset Statistics Total Size: ~87GB Number of Splits: 32 (split-0.tar.gz to split-31.tar.gz) Audio Format: WAV files Source: Speech edited from LibriLight… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/SINE.audioaudio-classificationn<1K0 likes90 downloads1y agoHugging Face28peaceAsh /floorplan-room-segmentation Floorplans Dataset This dataset is derived from the Floorplans Diff dataset and has been curated by removing all unannotated images to ensure clean and consistent training data. It is designed for semantic image segmentation, specifically focusing on identifying and segmenting rooms within floorplan images. Each sample consists of an image paired with a corresponding segmentation mask, enabling models to learn pixel-level classification for the room class. Origin This… See the full description on the dataset page: https://huggingface.co/datasets/peaceAsh/floorplan-room-segmentation.imageimage-segmentation1K<n<10K1 likes86 downloads10mo agoHugging Face29willxxy /ecg-comprehension-r-peak-count-789480-5000-250-2500text100K<n<1M0 likes80 downloads7mo agoHugging Face30PEARLS-Lab /TALES-Trajectories TALES Trajectories Agent trajectory data from the TALES: Text Adventure Learning Environment Suite benchmark. TALES: Text Adventure Learning Environment Suite Christopher Zhang Cui, Xingdi Yuan, Ziang Xiao, Prithviraj Ammanabrolu, Marc-Alexandre Côté arXiv:2504.14128 Links: Paper | GitHub Leaderboard Top agents ranked by average best normalized score per game across 122 games, each repeated over 5 seeds (610 total). Scores reflect the highest normalized score… See the full description on the dataset page: https://huggingface.co/datasets/PEARLS-Lab/TALES-Trajectories.tabulartext-generation10K<n<100K0 likes72 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.