datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.Palladium-1M-Preview
💎 Palladium-1M: High-Density Information for Efficient LLM Training
Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System.
Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise."
The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.PALATE
PALATE Dataset
PALATE contains de-identified human–role-playing-agent conversations,
satisfaction annotations, frozen session-level splits, bilingual character
cards, and the scoring rubrics used by the PALATE benchmark.
Related resources:
Code: Zhuyh1139/PALATE
Five user-simulator adapters:
muset-ai/PALATE-LoRA
The dataset stores source annotations rather than ready-to-train examples.
Use the processing command in the PALATE GitHub repository to construct
role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.bias-correction-palestine-protocol
Dataset Card for LLM Bias Correction (Palestine/Israel Context)
This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel.
Dataset Structure
The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.pall
PALL — Dental Training Corpus
Open training corpus for PALL-Text, a
dental-domain Llama-3.1-8B. Contains three subsets covering the full
CPT → SFT → DPO post-training pipeline.
Developed by: Harisundar R
License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms)
Language: English (with some multilingual medical Q&A)
Dataset structure
Subset
Schema
Train
Val
Total
cpt
{ "text", "source" }
401,900
4,059
405,959
sft
{… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.myanmar-english-pali-dictionary
Myanmar–English–Pali Dictionary
Dataset Summary
This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein).
It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary.
The dataset is intended for research and educational purposes, including but not limited to:
Natural Language Processing (NLP)
Machine Translation (MT)
Lexicography
Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.meridian-palace-training
🏨 The Meridian Palace — AI Hotel Staff Training Data
16,000 multi-turn conversations for fine-tuning a small LLM to act as 8 AI hotel staff roles at a luxury 5-star hotel.
Dataset Details
Train: 15,200 conversations
Validation: 800 conversations
Format: ChatML (system/user/assistant messages)
AI Roles Covered
Reservation Agent
Concierge
Guest Help Desk
Room Service
Virtual Front Desk
Cashier Assistant
Housekeeping Coordinator
Security Assistant… See the full description on the dataset page: https://huggingface.co/datasets/himu1780/meridian-palace-training.palladium-stem-preview-25k
⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus
"The Top 0.17% of the Open Web."
Overview
This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability.
The "Goldilocks" Methodology
Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.bigcodebench-plus
BCBPlus — BigCodeBench-Plus (Palaestra Curated)
A fixed fork of bubbleresearch/bigcodebench-plus with spec ambiguities, test bugs, and broken canonical solutions corrected.
Version: v1.0.2
Rows: 1136
SHA256: 3b05c95c55e018d5...
Upstream: bubbleresearch/bigcodebench-plus
Status breakdown
Status
Count
active
1136
Curation philosophy
Deterministic docstring examples are spec. Tests must agree with them.
Library conventions are binding. A test… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/bigcodebench-plus.ARIA-PaLM-text
