datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BrowseComp-ZH
🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese
BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.ucmo
UCMO — Non-Contaminated Math Olympiads
Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation.
Version: v0.0.4
Rows: 429
SHA256: 1f5f51a09ccd3674...
Stats
Answer type
Count
closed_form
121
numeric
170
open_ended
128
set
10
Total sources: 48
Schema
Each row:
Field
Description
id
Unique identifier (e.g., aime_i_2026_15)
source
Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.pallas_splitted_18cSecKnowledge-Eval
SecKnowledge 2.0 Evaluation Benchmark
The official evaluation benchmark suite from Toward Cybersecurity-Expert Small Language Models (ICML 2026), where we introduce the CyberPal 2.0 model family alongside these benchmarks. This repository releases the internal evaluation datasets developed to assess LLMs on core cybersecurity capabilities that existing public benchmarks do not adequately cover: adversarial robustness on CTI knowledge, cross-taxonomy reasoning, consequence-centric… See the full description on the dataset page: https://huggingface.co/datasets/cyber-pal-security/SecKnowledge-Eval.palette-bench-ko
PALETTE-BENCH-KO — Korean Enterprise Document Benchmark
Version: 0.1 (seed) · License: CC BY 4.0 · Language: Korean (ko)
What this is
The first public benchmark for Korean enterprise document work — the drafting,
extraction, and compliance tasks that office staff actually do, which existing Korean
benchmarks (KMMLU, HAE-RAE, KoBALT, LogicKor) do not cover. A landscape sweep (2026-08)
found no public benchmark testing 공문서/품의서 drafting, 회의록→결정 extraction, or… See the full description on the dataset page: https://huggingface.co/datasets/palette-lab/palette-bench-ko.plot-palette-100k
Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.palm
🏝️ Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs
🏆 Best Resource Paper Award - ACL 2025
Overview
Palm is the first comprehensive, human-created Arabic instruction dataset that is both culturally and linguistically diverse and inclusive. Created through a year-long community-driven effort by 44 researchers across 22 Arab countries, Palm represents a landmark achievement in Arabic NLP.
Key Features
🌍 All-Inclusive… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palm.palmx_2025_subtask1_culture🏷️ PalmX 2025 — General Culture Evaluation (PalmX-GC)
Dataset Summary
PalmX-GC evaluates a model’s grasp of general Arab culture—customs, history, geography, arts, cuisine, notable figures, and everyday life across the 22 Arab League countries.
Every item is written in Modern Standard Arabic (MSA). The dataset powers Subtask 1 of the PalmX 2025 shared task.
Dataset Structure
Split
# MCQs
Release Date
Notes
Train
2000
10 Jun 2025
With gold answers
Dev
500
10… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palmx_2025_subtask1_culture.pmt-licitacoes-qa-instruct
PMT Licitações QA — Dataset de Fine-Tuning
Dataset de perguntas e respostas sobre licitações públicas da Prefeitura Municipal de Teresina (PMT), estruturado no formato instrução-entrada-saída para fine-tuning de modelos de linguagem.
Descrição
Este dataset foi construído a partir de documentos de licitação disponibilizados publicamente no portal da Prefeitura Municipal de Teresina.
Os pares de QA foram gerados por meio de destilação de conhecimento utilizando o modelo… See the full description on the dataset page: https://huggingface.co/datasets/palaciodata/pmt-licitacoes-qa-instruct.pall
PALL — Dental Training Corpus
Open training corpus for PALL-Text, a
dental-domain Llama-3.1-8B. Contains three subsets covering the full
CPT → SFT → DPO post-training pipeline.
Developed by: Harisundar R
License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms)
Language: English (with some multilingual medical Q&A)
Dataset structure
Subset
Schema
Train
Val
Total
cpt
{ "text", "source" }
401,900
4,059
405,959
sft
{… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.pali-myanmar-dictionary-corpus
Pali-Myanmar Dictionary Corpus (Instruction-Ready)
Dataset Summary
The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning.
Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.arabic-palestinian-levantine-sample
4FACTORS — Palestinian Levantine Conversational Sample
50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data.
What this is
Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.GSM-PALitalian_dataset_mix
Dataset Card for Dataset Name
This dataset represents a collection of the most downloaded Italian datasets.
Dataset Details
Dataset Description
This dataset represents a collection of the most downloaded Italian datasets:
WasamiKirua/samantha-ita
mii-community/ultrafeedback-translated-ita
mchl-labs/stambecco_data_it
efederici/fisica
FreedomIntelligence/sharegpt-italian
Curated by: Enzo Palmisano
Language(s) (NLP): Italian
License: Apache 2.0
palmx_2025_subtask2_islamic
🏷️ PalmX 2025 — Islamic Culture Evaluation (PalmX-IC)
Dataset Summary
PalmX-IC assesses a model’s knowledge of Islamic culture—rituals, Qurʾān verses, Ḥadīth, historic events, jurisprudence, and religious holidays—core elements of life across the Arab world.All items are authored in Modern Standard Arabic (MSA) . The dataset powers Subtask 2 of the PalmX 2025 shared task.
Dataset Structure
Split
# MCQs
Release Date
Notes
Train
600
10 Jun 2025
With… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palmx_2025_subtask2_islamic.Palestinian_Truth_Englishe_commerce_customer_service_squadPalGeoLLM
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [Ahmad Budairi1, Belal Hamdeh1, and Mitri Khoury]
Funded by [optional]: [Birzeit University]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [Arabic]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/RuwaYafa/PalGeoLLM.
