datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
egodistillopendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.peacock-data-public-datasets-sangrahaCIDER
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Paper | Code
Dataset for the COLM 2026 paper CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
CIDER is a dataset of privacy disclosure decisions collected from real users. It consists of 14,850 annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios.
What can you do with… See the full description on the dataset page: https://huggingface.co/datasets/peach-lab/CIDER.conceptnet_en_simplepeak-anchor-content-35kArabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.peacock-data-public-datasetspeak-search-content-70kpeak-intent-50pear-data
🍐 PEAR MT Evaluation Data
Overview
This dataset contains the pairwise Machine Translation evaluation data used to
train and evaluate
PEAR: Pairwise Evaluation
for Automatic Relative Scoring in Machine Translation.
Each example contains:
a source segment;
an optional human reference;
two candidate translations;
the corresponding MT system identifiers;
human quality scores for both candidates;
contextual metadata such as year, language pair, and domain.
The… See the full description on the dataset page: https://huggingface.co/datasets/Prosho/pear-data.PEACE
PEACE: Empowering Geologic Map Holistic Understanding with MLLMs
[Code] [Paper] [Data]
Introduction
We construct a geologic map benchmark, GeoMap-Bench, to evaluate the performance of MLLMs on geologic map understanding across different abilities, the overview of it is as shown in below Table.
Property
Description
Source
USGS(English)
CGS(Chinese)
Content
Image-question pair… See the full description on the dataset page: https://huggingface.co/datasets/microsoft/PEACE.peak-search-300kpeanuts-opt-6.7b
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,457 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-opt-6.7b.2025_DCASE_AudioQA_Official
Audio SFT / Post-Training Data
The proposed audio question answering (AQA) dataset
with three categories: Bioacoustics QA (BQA), Temporal Soundscapes QA (TSQA), and Complex QA (CQA)
DCASE 2025 Task Description
Audio QA Model Baseline
Watkins Marine Mammal Sound Database
"Watkins Marine Mammal Sound Database, Woods Hole Oceanographic Institution and the New Bedford Whaling Museum."
📢 Post-Challenge Research Note
While the DCASE 2025 Challenge… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/2025_DCASE_AudioQA_Official.and-peacenmrexp-cnmr-peaklist-1.5Mmeow-tea-oolongNeko-v1private for working in progress. ACL 2025 underview.
pearl_benchmark
PEARL-Benchmark: A benchmark for evaluating phrase representations
Learning High-Quality and General-Purpose Phrase Representations.
Lihu Chen, Gaël Varoquaux, Fabian M. Suchanek.
Accepted by EACL Findings 2024
Our PEARL Benchmark contains 9 phrase-level datasets of five types of tasks, which cover both the field of data science and natural language processing.
Description
Paraphrase Classification: PPDB and PPDBfiltered (Wang et al., 2021)
Phrase Similarity:… See the full description on the dataset page: https://huggingface.co/datasets/Lihuchen/pearl_benchmark.conceptnet_en_nomalizedThis is the English part of the ConceptNet and we have removed the useless information.
StreamGaze_v2
StreamGaze Dataset
StreamGaze is a comprehensive streaming video benchmark for evaluating MLLMs on gaze-based QA tasks across past, present, and future contexts.
Companion dataset: The EgoGazeVQA dataset is hosted separately at Peanuttoad/gaze_dataset.
📁 Dataset Structure
streamgaze/
├── metadata/
│ ├── egtea.csv # EGTEA fixation metadata
│ ├── egoexolearn.csv # EgoExoLearn fixation metadata
│ └── holoassist.csv # HoloAssist… See the full description on the dataset page: https://huggingface.co/datasets/Peanuttoad/StreamGaze_v2.peak-anchor-40kpeak-text-with-context-2mPEARLpeacock-data-public-datasets-hubtcg-frame-removal-dataset
TCG Frame Removal Dataset
547 paired examples for training instruction-editing models that strip the frame, text,
and UI elements from trading-card images and extend the artwork to a seamless full-bleed
illustration. This is the training set for the
TCG Frame Removal LoRA (FLUX.2-Klein 4B) model
(weights).
Game
Pairs
Magic: The Gathering
304
Digimon
154
Pokémon
63
Yu-Gi-Oh!
26
Fields
id (string): unique card slug, prefixed by game (mtg-… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/tcg-frame-removal-dataset.PEaCEtajik-asr-corpus-v3
tajik-asr-corpus-v3
1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.
