datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
phonemizer-dicts
Phonemizer Dicts
Pre-generated IPA dictionaries for GPL-free text-to-phonemes lookup.
Files
en-us.tsv — 124K English (US) words, tab-separated word<TAB>IPA
Provenance
Generated by running espeak-ng over an English wordlist. The TSV is program output; espeak-ng source (GPL-3.0) is not redistributed here.
Regeneration
See scripts/generate-espeak-dict.py in the tts-rd-team repo.
Paladin_TCGA_CPTAC_omicsPaladin TCGA & CPTAC Spatial Omics Maps
Ready-to-use patch-level and slide-level spatial omics maps inferred by
Paladin from TCGA and CPTAC
whole-slide images. The released .Paladin.h5 files can be analyzed directly
without rerunning WSI inference.
The collection is populated in stages. Check Files and versions for the
cohorts currently available.
Spatial multi-omics example
The panels show the H&E WSI, a reference tumor mask, CNV burden, TP53 CNV,
DNA-methylation… See the full description on the dataset page: https://huggingface.co/datasets/zhihuanglab/Paladin_TCGA_CPTAC_omics.PalmDex
🤖 PalmDex: An Embodiment-Agnostic Tactile Manipulation Dataset
A growing, multimodal robotic manipulation dataset featuring synchronized dual-camera video, dual-hand tactile sensing, and hand pose tracking across diverse real-world environments. All demonstrations are collected via human teleoperation — without any specific robot embodiment — and annotated at the action-segment level with rich categorical labels.
[!NOTE]
This is a living dataset. New environments, tasks, and… See the full description on the dataset page: https://huggingface.co/datasets/Rimbot/PalmDex.BrowseComp-ZH
🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese
BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.persona_in_palpaloma
Dataset Card for Paloma
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.PALM
Dataset Card for PALM-XR20A (1000 Sample Subset)
This is a FiftyOne dataset with 1000 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
from huggingface_hub import snapshot_download
# Download the dataset snapshot to the current working directory
snapshot_download(
repo_id="Voxel51/PALM",
local_dir=".",
repo_type="dataset"
)
# Load dataset from current… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/PALM.Palmpad_Dataset
Palmpad Dataset
The Palmpad Dataset is a publicly available multimodal dataset designed for research in finger-to-palm recognition. This dataset contains data from users, divided into two groups (Group 1 and Group 2). Each user's data consists of multiple video files in .avi format, along with corresponding frame-level annotation files in .txt format.
Overview
Paper Title: Palmpad: Enabling Real-Time Index-to-Palm Touch Interaction with a Single RGB Camera
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Teburile/Palmpad_Dataset.CUHK-X_Small_Model_Track
CUHK-X — Small Model Track
Multimodal human action recognition (classification).
Given a multimodal clip, predict its action class (action_id, 0–39, 40 classes).
Repository layout
.
├── Training/
│ ├── class_mapping.csv # action_id <-> action_name (40 classes)
│ └── data/
│ └── HAR.z01 … HAR.z08 + HAR.zip # multi-volume zip
│ → HAR/data/<modality>/<action>/<user>/<trial>/<files>
└── Testing/
├── data/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Small_Model_Track.PALM
🩺 PALM — Pathologic Myopia Fundus Image Dataset
Image: Dataset Samples.
📘 Overview
PALM (Pathologic Myopia) is a publicly available fundus image dataset developed for detecting pathologic myopia (PM) and analyzing associated retinal lesions and anatomical structures.
It was released for the Pathologic Myopia Challenge (PALM), hosted by the Chinese Academy of Sciences and Sun Yat-sen University, and… See the full description on the dataset page: https://huggingface.co/datasets/ctmedtech/PALM.pal-resultscoco-paligemmaforce-prompting-dataset-creationro-real-estate-listings
Dataset Card for Romanian Real Estate Listings Dataset
Dataset Details
Dataset Description
This dataset contains publicly available real estate listings scraped from Romanian property websites.It includes structured information such as price, location, number of rooms, and surface area.The dataset is continuously updated through an automated data pipeline.
Curated by: Flavius Paler
Language(s): Romanian
License: MIT
Uses… See the full description on the dataset page: https://huggingface.co/datasets/palerflavius/ro-real-estate-listings.attackdex-paldeaSingle pokemon datasets containing all the attacks (from levelling or TMs) learnable by the relative monster. All the data refer to the Paldea region and they come from the project discussed in https://medium.com/@virtualmartire/i-built-an-algorithm-that-finds-the-optimal-pokemon-team-01ea152824a9.
CUHK-X_Large_Model_Track
CUHK-X — Large Model Track
Multimodal video question answering (VQA) over human daily-activity clips recorded at home.
Given a short multimodal video, answer multiple-choice questions about it.
Repository layout
.
├── Training/
│ ├── training_qa.csv # questions + answers
│ ├── modality_list.csv # which modalities each clip has
│ └── data/
│ ├── HARn.zip → HARn/<action>/<user>/<trial>/<modality>/<modality>.mp4
│ └── HAU.zip →… See the full description on the dataset page: https://huggingface.co/datasets/Kevin-Pal/CUHK-X_Large_Model_Track.xBD_Dataset_minepali_processed_983814details_Mihaiii__Pallas-0.4
Dataset Card for Evaluation run of Mihaiii/Pallas-0.4
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.4 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.4.details_Mihaiii__Pallas-0.3
Dataset Card for Evaluation run of Mihaiii/Pallas-0.3
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.3 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.3.details_Writer__palmyra-med-20b
Dataset Card for Evaluation run of Writer/palmyra-med-20b
Dataset Summary
Dataset automatically created during the evaluation run of model Writer/palmyra-med-20b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Writer__palmyra-med-20b.controlnet-color-palette-20K
ControlNet Color Palette Dataset
This dataset contains resized images (512×512) and pixelated conditioning maps
generated using a 8×8 grid.
details_Mihaiii__Pallas-0.2
Dataset Card for Evaluation run of Mihaiii/Pallas-0.2
Dataset Summary
Dataset automatically created during the evaluation run of model Mihaiii/Pallas-0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Mihaiii__Pallas-0.2.ASVspoof2021_DF
ASVspoof 2021 DF
Benchmark-ready packaging of the DeepFake (DF) evaluation partition from ASVspoof 2021 for speech anti-spoofing and synthetic / deepfake voice detection.
Overview
This dataset contains the DF evaluation subset of the ASVspoof 2021 challenge. The task is binary classification: bonafide (genuine human speech) vs. spoof (synthetic, converted, or otherwise manipulated speech). The original dataset is available at… See the full description on the dataset page: https://huggingface.co/datasets/palvitha06/ASVspoof2021_DF.ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.pali-tripitaka-thai-script-siamrath-version
Multi-File CSV Dataset
คำอธิบาย
พระไตรปิฎกภาษาบาลี อักษรไทยฉบับสยามรัฏฐ จำนวน ๔๕ เล่ม
ชุดข้อมูลนี้ประกอบด้วยไฟล์ CSV หลายไฟล์
01/010001.csv: เล่ม 1 หน้า 1
01/010002.csv: เล่ม 1 หน้า 2
...
02/020001.csv: เล่ม 2 หน้า 1
...
คำอธิบายของแต่ละเล่ม
เล่ม ๑: วินย. มหาวิภงฺโค (๑)
เล่ม ๒: วินย. มหาวิภงฺโค (๒)
เล่ม ๓: วินย. ภิกฺขุนีวิภงฺโค
เล่ม ๔: วินย. มหาวคฺโค (๑)
เล่ม ๕: วินย. มหาวคฺโค (๒)
เล่ม ๖: วินย. จุลฺลวคฺโค (๑)
เล่ม ๗: วินย. จุลฺลวคฺโค (๒)
เล่ม ๘: วินย.… See the full description on the dataset page: https://huggingface.co/datasets/uisp/pali-tripitaka-thai-script-siamrath-version.controlnet-color-palette-20K-e1paloma_validationpaloma_programming_languagesDiscordLog
