datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
paloma
Dataset Card for Paloma
Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.PALL-VLM-data
PALL-VLM-data — Dental Vision-Language Dataset
The training dataset for Harisundar/PALL-VLM,
a dental vision-language model. It contains 32,884 records over 52,461 images,
formatted as image+text conversations for LLaVA-style instruction tuning.
Curated by: Harisundar R
Used by: Harisundar/PALL-VLM · PALL on GitHub
Language: English
Layout
vlm_train/
├── images/ # 52,461 dental images
├── train.jsonl # 29,667 records
├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.palestine
Palestine Dataset 🇵🇸
A curated dataset focused on authentic Palestinian history, narratives and reporting.
Data Sources 📊
decolonizepalestine.com - Educational content and historical documentation
electronicintifada.net - Hundreds of articles - news, analysis, and more
palianswers.com - A crowdsourced database of short responses to Zionist claims
english.khamenei.ir - Articles related to Palestine
mondoweiss.net - Hundreds of articles - news, analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/mlibre/palestine.PaladinDatacolor-palettes-sdpallatom-ligand-assets
LevinHarness/pallatom-ligand-assets — public mirror of third-party runtime assets
This dataset is a public mirror of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/pallatom-ligand-assets.Palladium-1M-Preview
💎 Palladium-1M: High-Density Information for Efficient LLM Training
Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System.
Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise."
The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.open-computer-using-agent
Dataset Description
This dataset is associated with the ongoing 'Nous Project' - creating a computer using agent based on open source models. The data was collected using Anthropic Claude 3.5 Sonnet Latest to record conversational state along with computer state data including:
Cursor position
Active windows
Computer display dimensions
System state information
User interactions
While the original interview problem covered only button clicks, this dataset is more comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/palkarpratik84/open-computer-using-agent.pall
PALL — Dental Training Corpus
Open training corpus for PALL-Text, a
dental-domain Llama-3.1-8B. Contains three subsets covering the full
CPT → SFT → DPO post-training pipeline.
Developed by: Harisundar R
License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms)
Language: English (with some multilingual medical Q&A)
Dataset structure
Subset
Schema
Train
Val
Total
cpt
{ "text", "source" }
401,900
4,059
405,959
sft
{… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
cinematic-mood-palette
Cinematic Mood Palette
Curated mappings between affective states and cinematic visual expression. The goal is to describe how filmmakers translate psychological affect into color and perceptual parameters.
~80 mappings, including emotional states, cinematic aesthetics, and spatial calibration points.
What This Is
A collection of anchor points in a 5-dimensional emotional space, each paired with corresponding cinematic color and perceptual parameters.
It functions… See the full description on the dataset page: https://huggingface.co/datasets/danielritchie/cinematic-mood-palette.PAlign-PAPI-personality_prompt.json-cleanedAdapted from
"Personality Alignment of Large Language Models" by Minjun Zhu and Linyi Yang and Yue Zhang
and the associated GitHub repository zhu-minjun/PAlign.
The contents of said repo were declared public domain; in that spirit, this Alpaca-formatted file has also been released as public domain.
myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.meridian-palace-training
🏨 The Meridian Palace — AI Hotel Staff Training Data
16,000 multi-turn conversations for fine-tuning a small LLM to act as 8 AI hotel staff roles at a luxury 5-star hotel.
Dataset Details
Train: 15,200 conversations
Validation: 800 conversations
Format: ChatML (system/user/assistant messages)
AI Roles Covered
Reservation Agent
Concierge
Guest Help Desk
Room Service
Virtual Front Desk
Cashier Assistant
Housekeeping Coordinator
Security Assistant… See the full description on the dataset page: https://huggingface.co/datasets/himu1780/meridian-palace-training.palladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.spider_SQL_PALM_PromptDataset for creating prompts for fine-tuning on Spider Dataset with Foreign and Primary Key Information as well as Schema information.
Palma-1.0
Palma-1.0 Dataset
Comprehensive Database of Global Palm Species
Palma-1.0 is a comprehensive exploration of palm species, including the PalmTraits 1.0 dataset enriched with data from GBIF, iNaturalist, Wikimedia Commons, and Plants of the World Online (POWO).
Dataset Overview
Palma-1.0 contains comprehensive data on 2,557 palm species across 181 genera. The dataset combines morphological traits, taxonomic classification, geographic distribution, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/kitsuiwebster/Palma-1.0.bigcodebench-plus
BCBPlus — BigCodeBench-Plus (Palaestra Curated)
A fixed fork of bubbleresearch/bigcodebench-plus with spec ambiguities, test bugs, and broken canonical solutions corrected.
Version: v1.0.2
Rows: 1136
SHA256: 3b05c95c55e018d5...
Upstream: bubbleresearch/bigcodebench-plus
Status breakdown
Status
Count
active
1136
Curation philosophy
Deterministic docstring examples are spec. Tests must agree with them.
Library conventions are binding. A test… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/bigcodebench-plus.Palette_nlu_service
Persian Sales NLU (seed)
Synthetic + templated Persian sales dialogues for NLU training.
Splits: train.jsonl, val.jsonl, test.jsonl
Each line: text, tokens, intent, slots (BIO), keep_mask, search_query.
P-ALIGNpallatom-ligand-assets
LevinHarness/pallatom-ligand-assets — public mirror of third-party runtime assets
This dataset is a public mirror of third-party runtime
assets required by the Levin Harness plugin(s) listed below. It is not an
official distribution: nothing here is published under this account's own
terms, and it is not affiliated with or endorsed by any upstream project.
Ownership and licensing
Every file remains the property of its upstream authors.
Each file keeps its… See the full description on the dataset page: https://huggingface.co/datasets/sgetttt/pallatom-ligand-assets.pale-madlad-data
license: mit
PaLe-MADLAD Data
Data used for training the PaLe-MADLAD model to translate from Proper Karelian, Livvi, Ludian, and Veps to Russian and vice versa. Every dataset entry represents a single text and comes as a list of sentences supplemented (where possible) with a list of translations into Russian. Our sources include:
VepKar: various articles, Biblical texts, folklore, and more in Proper Karelian, Livvi, Ludian, and Veps, mostly translated into Russian… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/pale-madlad-data.PALALDINvinaya-pitaka-pali-myanmar-parallel
Vinaya Pitaka: Pali-Myanmar Parallel Dataset
Description
This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation.
The dataset covers all five major volumes of the Vinaya:
Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်)
Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်)
Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ)
Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.paloalma__ECE-TW3-JRGL-V1-details
Dataset Card for Evaluation run of paloalma/ECE-TW3-JRGL-V1
Dataset automatically created during the evaluation run of model paloalma/ECE-TW3-JRGL-V1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__ECE-TW3-JRGL-V1-details.pali-viet
