datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
BrowseComp-ZH
🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese
BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.persona_in_palpaloma
Dataset Card for Paloma
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.Palmpad_Dataset
Palmpad Dataset
The Palmpad Dataset is a publicly available multimodal dataset designed for research in finger-to-palm recognition. This dataset contains data from users, divided into two groups (Group 1 and Group 2). Each user's data consists of multiple video files in .avi format, along with corresponding frame-level annotation files in .txt format.
Overview
Paper Title: Palmpad: Enabling Real-Time Index-to-Palm Touch Interaction with a Single RGB Camera
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Teburile/Palmpad_Dataset.CUHK-X_Small_Model_Track
CUHK-X — Small Model Track
Multimodal human action recognition (classification).
Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes).
Repository layout
.
├── Training/
│ ├── class_mapping.csv # action_id <-> action_name (40 classes)
│ └── data/
│ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip
│ → HAR/data/<modality>/<action>/<user>/<trial>/<files>
└── Testing/
├── data/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.coco-paligemmaro-real-estate-listings
Dataset Card for Romanian Real Estate Listings Dataset
Dataset Details
Dataset Description
This dataset contains publicly available real estate listings scraped from Romanian property websites.It includes structured information such as price, location, number of rooms, and surface area.The dataset is continuously updated through an automated data pipeline.
Curated by: Flavius Paler
Language(s): Romanian
License: MIT
Uses… See the full description on the dataset page: https://huggingface.co/datasets/palerflavius/ro-real-estate-listings.attackdex-paldeaSingle pokemon datasets containing all the attacks (from levelling or TMs) learnable by the relative monster. All the data refer to the Paldea region and they come from the project discussed in https://medium.com/@virtualmartire/i-built-an-algorithm-that-finds-the-optimal-pokemon-team-01ea152824a9.
pali_processed_983814controlnet-color-palette-20K
ControlNet Color Palette Dataset
This dataset contains resized images (512×512) and pixelated conditioning maps
generated using a 8×8 grid.
ASVspoof2021_DF
ASVspoof 2021 DF
Benchmark-ready packaging of the DeepFake (DF) evaluation partition from ASVspoof 2021 for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
This dataset contains the DF evaluation subset of the ASVspoof 2021 challenge. The task is binary classification: bonafide (genuine human speech) vs. spoof (synthetic, converted, or otherwise manipulated speech). The original dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/palvitha06/ASVspoof2021_DF.ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.controlnet-color-palette-20K-e1paloma_validationpaloma_programming_languagescontrolnet-color-palette-7K
ControlNet Color Palette Dataset
This dataset contains resized images (768×768) and pixelated conditioning maps
generated using a 8×8 grid.
africa-synth-cerebral-palsy-synthetic-dataset-all
African Cerebral Palsy Synthetic Dataset | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: csv - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-cerebral-palsy-synthetic-dataset-all.paloma_subredditspali-tripitaka-thai-script-siamrath-version
📚 พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฐ (๔๕ เล่ม)
ต้นฉบับสามารถเข้าถึงได้ที่: 84000 พระธรรมขันธ์ 84000.org พระไตรปิฎก
🧾 รายการพระไตรปิฎก
📘 เล่ม ๑–๘: วินัยปิฎก
เล่ม ๑: มหาวิภงฺโค (๑)
เล่ม ๒: มหาวิภงฺโค (๒)
เล่ม ๓: ภิกฺขุนีวิภงฺโค
เล่ม ๔: มหาวคฺโค (๑)
เล่ม ๕: มหาวคฺโค (๒)
เล่ม ๖: จุลฺลวคฺโค (๑)
เล่ม ๗: จุลฺลวคฺโค (๒)
เล่ม ๘: ปริวาโร
📗 เล่ม ๙–๒๕: สุตตันตปิฎก
เล่ม ๙–๑๑: ทีฆนิกาย
เล่ม ๑๒–๑๔: มัชฌิมนิกาย
เล่ม ๑๕–๑๙: สังยุตตนิกาย… See the full description on the dataset page: https://huggingface.co/datasets/mgprogm/pali-tripitaka-thai-script-siamrath-version.iChallenge-PALM19
iChallenge-PALM19 — PALM: PAthoLogic Myopia
1200 color fundus photographs from 720 myopia-clinic subjects (left eyes
only) at Zhongshan Ophthalmic Center, Sun Yat-sen University — the PALM dataset
from the ISBI 2019 iChallenge satellite event, with pixel-level annotation of
three targets.
Modality: color fundus photography (2D RGB)
Organ: retina / eye
Resolutions: 2124×2056 (Zeiss Visucam 500, n=1047) · 1444×1444 (Canon CR-2, n=153)
Targets: optic disc · patchy retinal atrophy… See the full description on the dataset page: https://huggingface.co/datasets/MedOtter/iChallenge-PALM19.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.palestine
Palestine Dataset 🇵🇸
A curated dataset focused on authentic Palestinian history, narratives and reporting.
Data Sources 📊
decolonizepalestine.com - Educational content and historical documentation
electronicintifada.net - Hundreds of articles - news, analysis, and more
palianswers.com - A crowdsourced database of short responses to Zionist claims
english.khamenei.ir - Articles related to Palestine
mondoweiss.net - Hundreds of articles - news, analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/mlibre/palestine.PALACE_GRPOReasoning-DeepSeek-R1-Distilled-1.4M-Alpaca-V2auto-pale
Dataset card for pale
Dataset summary
This dataset contains league of legends champions' quotes parsed from fandom.
See dataset usage example at google colab.
The dataset is available in the following configurations:
vanilla - all data pulled from the website without significant modifications apart from the web page structure parsing;
quotes - truncated version of the corpus, which does't contain sound effects;
annotated - an extended version of the full configuration… See the full description on the dataset page: https://huggingface.co/datasets/zeio/auto-pale.palmer-penguins
Palmer Penguins
The Palmer penguins dataset by Allison Horst, Alison Hill, and Kristen Gorman was first made publicly available as an R package.
The goal of the Palmer Penguins dataset is to replace the highly overused Iris dataset for data exploration & visualization.
However, now you can use Palmer penguins on huggingface!
License
Data are available by CC-0 license in accordance with the Palmer Station LTER Data Policy and the LTER Data Access Policy for Type I data.… See the full description on the dataset page: https://huggingface.co/datasets/SIH/palmer-penguins.vid-guard-rlhf-unsafecontrolnet-color-palette-5K_OG
ControlNet Color Palette Dataset
This dataset contains resized images (768×768) and pixelated conditioning maps
generated using a 8×8 grid.
