datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.auto-pale
Dataset card for pale
Dataset summary
This dataset contains league of legends champions' quotes parsed from fandom.
See dataset usage example at google colab.
The dataset is available in the following configurations:
vanilla - all data pulled from the website without significant modifications apart from the web page structure parsing;
quotes - truncated version of the corpus, which does't contain sound effects;
annotated - an extended version of the full configuration… See the full description on the dataset page: https://huggingface.co/datasets/zeio/auto-pale.pallas_splitted_18cpaleo-hebrew-seals-synthetic
PaleoHebrew-Seals Synthetic Corpus
This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions.
Why this dataset is needed
Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level.
Overview
The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.task850_synthetic_longest_palindrome
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task850_synthetic_longest_palindrome
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task850_synthetic_longest_palindrome.Palladium-1M-Preview
💎 Palladium-1M: High-Density Information for Efficient LLM Training
Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System.
Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise."
The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.palm
🏝️ Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
🏆 Best Resource Paper Award - ACL 2025
Overview
Palm is the first comprehensive, human-created Arabic instruction dataset that is both culturally and linguistically diverse and inclusive. Created through a year-long community-driven effort by 44 researchers across 22 Arab countries, Palm represents a landmark achievement in Arabic NLP.
Key Features
🌍 All-Inclusive… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palm.paleo-hebrew-seals-unambiguous
PaleoHebrew-Seals Real Benchmark (Unambiguous Subset)
This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs.
Why this dataset is needed
Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.pall
PALL — Dental Training Corpus
Open training corpus for PALL-Text, a
dental-domain Llama-3.1-8B. Contains three subsets covering the full
CPT → SFT → DPO post-training pipeline.
Developed by: Harisundar R
License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms)
Language: English (with some multilingual medical Q&A)
Dataset structure
Subset
Schema
Train
Val
Total
cpt
{ "text", "source" }
401,900
4,059
405,959
sft
{… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.meridian-palace-training
🏨 The Meridian Palace — AI Hotel Staff Training Data
16,000 multi-turn conversations for fine-tuning a small LLM to act as 8 AI hotel staff roles at a luxury 5-star hotel.
Dataset Details
Train: 15,200 conversations
Validation: 800 conversations
Format: ChatML (system/user/assistant messages)
AI Roles Covered
Reservation Agent
Concierge
Guest Help Desk
Room Service
Virtual Front Desk
Cashier Assistant
Housekeeping Coordinator
Security Assistant… See the full description on the dataset page: https://huggingface.co/datasets/himu1780/meridian-palace-training.palladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.bigcodebench-plus
BCBPlus — BigCodeBench-Plus (Palaestra Curated)
A fixed fork of bubbleresearch/bigcodebench-plus with spec ambiguities, test bugs, and broken canonical solutions corrected.
Version: v1.0.2
Rows: 1136
SHA256: 3b05c95c55e018d5...
Upstream: bubbleresearch/bigcodebench-plus
Status breakdown
Status
Count
active
1136
Curation philosophy
Deterministic docstring examples are spec. Tests must agree with them.
Library conventions are binding. A test… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/bigcodebench-plus.italian_dataset_mix
Dataset Card for Dataset Name
This dataset represents a collection of the most downloaded Italian datasets.
Dataset Details
Dataset Description
This dataset represents a collection of the most downloaded Italian datasets:
WasamiKirua/samantha-ita
mii-community/ultrafeedback-translated-ita
mchl-labs/stambecco_data_it
efederici/fisica
FreedomIntelligence/sharegpt-italian
Curated by: Enzo Palmisano
Language(s) (NLP): Italian
License: Apache 2.0
pallasbench-robust-gpu-a100
PallasBench: Robust Pallas GPU Kernel Benchmark (A100)
39/45 kernels passing on NVIDIA A100 80GB -- the first GPU-focused evaluation of JAX Pallas kernels.
What is this?
PallasBench is a suite of 45 JAX Pallas kernels across 3 difficulty levels. The original kernels were designed for TPU and failed on GPU because Pallas compiles to Triton on NVIDIA hardware, which has strict block size limits that TPU's Mosaic compiler does not.
We fixed all 45 kernels for GPU… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/pallasbench-robust-gpu-a100.Israel-palestine-war
Dataset Card for "Israel-palestine-war"
This Demo dataset is related to the research paper entitle "Online News Channel Streaming: A Comprehensive Analysis of Channel and User Engagement during the Israel-Palestine Conflict".
PREPRINT (Version 1) available at Research Square https://www.researchsquare.com/article/rs-3927576/latest
User Comments on News YouTube channels During Current War of Palstine & Israel Oct-2023.
Demo dataset size: {'NBCNews': 188490, 'aljazeeraenglish':… See the full description on the dataset page: https://huggingface.co/datasets/alsubari/Israel-palestine-war.task372_synthetic_palindrome_numbers
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task372_synthetic_palindrome_numbers
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task372_synthetic_palindrome_numbers.ARIA-PaLM-textpalestinian-cultural-knowledge
Palestinian Cultural Knowledge Corpus
v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload
(484 Arabic Wikipedia documents only). This release expands to the full 5-source
corpus below and moves the data to data/full_corpus/.
A multi-source Arabic/English text corpus about Palestinian history, culture, and
heritage, built for the Palestinian Cultural Knowledge
Platform
— a RAG + knowledge-graph research project. 882 documents, ~890K words, collected
and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.
