datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
opendatalab-experimental-nmr-peaks
OpenDataLab Experimental NMR Peaks Dataset
Dataset Description
This dataset contains experimental NMR (Nuclear Magnetic Resonance) peak sequences extracted from the OpenDataLab experimental spectra database. The dataset includes both H-NMR and C-NMR peak sequences for chemical compounds, along with their SMILES representations and molecular formulas.
Dataset Summary
Total Samples: 533,595 compounds
Batches: 333 batch files
Data Source: Experimental spectra… See the full description on the dataset page: https://huggingface.co/datasets/snehasis19/opendatalab-experimental-nmr-peaks.common-voice-scripted-speech-26
Common Voice Scripted Speech
A row-normalized multilingual ASR dataset built from Mozilla Data Collective
Common Voice Scripted Speech. Each upstream archive is converted to appendable
parquet shards under data/<upstream_split>/, one shard per source archive and
split, with audio bytes embedded in an audio struct column.
Status
Manifest languages: 60
Languages uploaded: 18
Columns
audio (bytes, path)
sentence, locale, language, upstream_split… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/common-voice-scripted-speech-26.conceptnet_en_simplepeak-anchor-content-35kArabic-VLM-Full-Pearl
💎 The Arabic VLM Dataset (Full Pearl Edition)
This repository contains the full, unreviewed dataset comprising 309K multimodal examples. This data was generated automatically using the agentic pipeline developed for the Pearl project, as described in our paper.
Disclaimer: This is the raw, synthetic data that has not been subject to human review. It was generated as part of the data creation process and is released for research purposes. It may contain noise, errors, or… See the full description on the dataset page: https://huggingface.co/datasets/MohamedRashad/Arabic-VLM-Full-Pearl.peak-search-content-70kpeak-intent-50pear-data
🍐 PEAR MT Evaluation Data
Overview
This dataset contains the pairwise Machine Translation evaluation data used to
train and evaluate
PEAR: Pairwise Evaluation
for Automatic Relative Scoring in Machine Translation.
Each example contains:
a source segment;
an optional human reference;
two candidate translations;
the corresponding MT system identifiers;
human quality scores for both candidates;
contextual metadata such as year, language pair, and domain.
The… See the full description on the dataset page: https://huggingface.co/datasets/Prosho/pear-data.peak-search-300kpeanuts-opt-6.7b
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,457 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-opt-6.7b.and-peacenmrexp-cnmr-peaklist-1.5MNeko-v1private for working in progress. ACL 2025 underview.
conceptnet_en_nomalizedThis is the English part of the ConceptNet and we have removed the useless information.
peak-anchor-40kpeak-text-with-context-2mpeacock-data-public-datasets-hubtcg-frame-removal-dataset
TCG Frame Removal Dataset
547 paired examples for training instruction-editing models that strip the frame, text,
and UI elements from trading-card images and extend the artwork to a seamless full-bleed
illustration. This is the training set for the
TCG Frame Removal LoRA (FLUX.2-Klein 4B) model
(weights).
Game
Pairs
Magic: The Gathering
304
Digimon
154
Pokémon
63
Yu-Gi-Oh!
26
Fields
id (string): unique card slug, prefixed by game (mtg-… See the full description on the dataset page: https://huggingface.co/datasets/pearsonkyle/tcg-frame-removal-dataset.PEaCEtajik-asr-corpus-v3
tajik-asr-corpus-v3
1,071 hours of Tajik ASR training data: 41 Tajik YouTube channels (~1,059 h, machine-labeled)
plus FLEURS tg_tj (11.8 h, gold). This is the dataset behind
Peacockery/omni-ctc-300m-tajik
(16.9% WER on FLEURS test, 37.6% on held-out conversational speech).
Layout
Hive-partitioned parquet under version=0/corpus=<source>/split=<split>/language=tgk_Cyrl/.
Each row holds text (the normalized label), audio_bytes (16 kHz mono FLAC as an int8 list),
and… See the full description on the dataset page: https://huggingface.co/datasets/Peacockery/tajik-asr-corpus-v3.HyPoradise-pilot
Dataset Name: Pilot dataset for Multi-domain ASR corrections
Description
This dataset is a pilot version of a larger dataset for automatic speech recognition (ASR) corrections across multiple domains.
It contains paired hypotheses and corrected transcriptions for various ASR tasks consolidated from PeacefulData/HyPoradise-v0
Structure
Data Split
The dataset is divided into training and test splits:
Training Data: 281,082 entries
Approximately… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/HyPoradise-pilot.peabody-testPEaCE-512pxecg-comprehension-r-peak-count-789480-10000-250-2500peanuts-flan-t5-xl
Peanut Comic Strip Dataset (Snoopy & Co.)
This is a dataset Peanuts comic strips from 1950/10/02 to 2000/02/13.
There are 77,456 panels extracted from 17,816 comic strips.
The dataset size is approximately 4.4G.
Each row in the dataset contains the following fields:
image: PIL.Image containing the extracted panel.
panel_name: unique identifier for the row.
characters: tuple[str, ...] of characters included in the comic strip the panel is part of.
themes: tuple[str, ...] of theme… See the full description on the dataset page: https://huggingface.co/datasets/afmck/peanuts-flan-t5-xl.pear-pests-dataset
🍐 Pear Leaf Pest & Disease Detection Dataset
This dataset supports object detection for pear leaf health monitoring: locating and classifying an insect pest and fungal/bacterial disease lesions directly on pear leaf images. It targets the timely detection and localization of foliar pear pests and diseases central to precision agriculture, where manual agronomist inspection is labor-intensive, subjective, and hard to scale.
The dataset consists of 2,210 annotated images (1,542… See the full description on the dataset page: https://huggingface.co/datasets/salahkhenfer/pear-pests-dataset.SINE
SINE Dataset
Overview
The Speech INfilling Edit (SINE) dataset is a comprehensive collection for speech deepfake detection and audio authenticity verification. This dataset contains ~87GB of audio data distributed across 32 splits, featuring both authentic and synthetically manipulated speech samples.
Dataset Statistics
Total Size: ~87GB
Number of Splits: 32 (split-0.tar.gz to split-31.tar.gz)
Audio Format: WAV files
Source: Speech edited from LibriLight… See the full description on the dataset page: https://huggingface.co/datasets/PeacefulData/SINE.floorplan-room-segmentation
Floorplans Dataset
This dataset is derived from the Floorplans Diff dataset and has been curated by removing all unannotated images to ensure clean and consistent training data.
It is designed for semantic image segmentation, specifically focusing on identifying and segmenting rooms within floorplan images.
Each sample consists of an image paired with a corresponding segmentation mask, enabling models to learn pixel-level classification for the room class.
Origin
This… See the full description on the dataset page: https://huggingface.co/datasets/peaceAsh/floorplan-room-segmentation.ecg-comprehension-r-peak-count-789480-5000-250-2500TALES-Trajectories
TALES Trajectories
Agent trajectory data from the TALES: Text Adventure Learning Environment Suite benchmark.
TALES: Text Adventure Learning Environment Suite
Christopher Zhang Cui, Xingdi Yuan, Ziang Xiao, Prithviraj Ammanabrolu, Marc-Alexandre Côté
arXiv:2504.14128
Links: Paper | GitHub
Leaderboard
Top agents ranked by average best normalized score per game across 122 games, each repeated over 5 seeds (610 total). Scores reflect the highest normalized score… See the full description on the dataset page: https://huggingface.co/datasets/PEARLS-Lab/TALES-Trajectories.
