CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PALIN2018 /BrowseComp-ZH 🧭 BrowseComp-ZH: Benchmarking the Web Browsing Ability of Large Language Models in Chinese BrowseComp-ZH is the first high-difficulty benchmark specifically designed to evaluate the real-world web browsing and reasoning capabilities of large language models (LLMs) in the Chinese information ecosystem. Inspired by BrowseComp (Wei et al., 2025), BrowseComp-ZH targets the unique linguistic, structural, and retrieval challenges of the Chinese web, including fragmented platforms… See the full description on the dataset page: https://huggingface.co/datasets/PALIN2018/BrowseComp-ZH.textquestion-answeringn<1K7 likes2.8k downloads1y agoHugging Face02palaestraresearch /ucmo UCMO — Non-Contaminated Math Olympiads Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation. Version: v0.0.4 Rows: 429 SHA256: 1f5f51a09ccd3674... Stats Answer type Count closed_form 121 numeric 170 open_ended 128 set 10 Total sources: 48 Schema Each row: Field Description id Unique identifier (e.g., aime_i_2026_15) source Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.textquestion-answeringn<1K0 likes295 downloads4mo agoHugging Face03FinchResearch /pallas_splitted_18ctexttext-classification1M<n<10M0 likes150 downloads3y agoHugging Face04cyber-pal-security /SecKnowledge-Eval SecKnowledge 2.0 Evaluation Benchmark The official evaluation benchmark suite from Toward Cybersecurity-Expert Small Language Models (ICML 2026), where we introduce the CyberPal 2.0 model family alongside these benchmarks. This repository releases the internal evaluation datasets developed to assess LLMs on core cybersecurity capabilities that existing public benchmarks do not adequately cover: adversarial robustness on CTI knowledge, cross-taxonomy reasoning, consequence-centric… See the full description on the dataset page: https://huggingface.co/datasets/cyber-pal-security/SecKnowledge-Eval.textquestion-answering1K<n<10K2 likes136 downloads4mo agoHugging Face05palette-lab /palette-bench-ko PALETTE-BENCH-KO — Korean Enterprise Document Benchmark Version: 0.1 (seed) · License: CC BY 4.0 · Language: Korean (ko) What this is The first public benchmark for Korean enterprise document work — the drafting, extraction, and compliance tasks that office staff actually do, which existing Korean benchmarks (KMMLU, HAE-RAE, KoBALT, LogicKor) do not cover. A landscape sweep (2026-08) found no public benchmark testing 공문서/품의서 drafting, 회의록→결정 extraction, or… See the full description on the dataset page: https://huggingface.co/datasets/palette-lab/palette-bench-ko.text-generationn<1K0 likes95 downloads1mo agoHugging Face06Hatman /plot-palette-100k Empowering Writers with a Universe of Ideas Plot Palette DataSet HuggingFace » Plot Palette was created to fine-tune large language models for creative writing, generating diverse outputs through iterative loops and seed data. It is designed to be run on a Linux system with systemctl for managing services. Included is the service structure, specific category prompts and ~100k data entries. The dataset is available here or… See the full description on the dataset page: https://huggingface.co/datasets/Hatman/plot-palette-100k.textquestion-answering10K<n<100K4 likes93 downloads2y agoHugging Face07UBC-NLP /palmgated 🏝️ Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs 🏆 Best Resource Paper Award - ACL 2025 Overview Palm is the first comprehensive, human-created Arabic instruction dataset that is both culturally and linguistically diverse and inclusive. Created through a year-long community-driven effort by 44 researchers across 22 Arab countries, Palm represents a landmark achievement in Arabic NLP. Key Features 🌍 All-Inclusive… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palm.textquestion-answering10K<n<100K21 likes71 downloads11mo agoHugging Face08UBC-NLP /palmx_2025_subtask1_culturegated🏷️ PalmX 2025 — General Culture Evaluation (PalmX-GC) Dataset Summary PalmX-GC evaluates a model’s grasp of general Arab culture—customs, history, geography, arts, cuisine, notable figures, and everyday life across the 22 Arab League countries. Every item is written in Modern Standard Arabic (MSA). The dataset powers Subtask 1 of the PalmX 2025 shared task. Dataset Structure Split # MCQs Release Date Notes Train 2000 10 Jun 2025 With gold answers Dev 500 10… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palmx_2025_subtask1_culture.textquestion-answering1K<n<10K1 likes68 downloads1y agoHugging Face09palaciodata /pmt-licitacoes-qa-instruct PMT Licitações QA — Dataset de Fine-Tuning Dataset de perguntas e respostas sobre licitações públicas da Prefeitura Municipal de Teresina (PMT), estruturado no formato instrução-entrada-saída para fine-tuning de modelos de linguagem. Descrição Este dataset foi construído a partir de documentos de licitação disponibilizados publicamente no portal da Prefeitura Municipal de Teresina. Os pares de QA foram gerados por meio de destilação de conhecimento utilizando o modelo… See the full description on the dataset page: https://huggingface.co/datasets/palaciodata/pmt-licitacoes-qa-instruct.textquestion-answering10K<n<100K0 likes56 downloads4mo agoHugging Face10Harisundar /pall PALL — Dental Training Corpus Open training corpus for PALL-Text, a dental-domain Llama-3.1-8B. Contains three subsets covering the full CPT → SFT → DPO post-training pipeline. Developed by: Harisundar R License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms) Language: English (with some multilingual medical Q&A) Dataset structure Subset Schema Train Val Total cpt { "text", "source" } 401,900 4,059 405,959 sft {… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.texttext-generation100K<n<1M1 likes51 downloads3mo agoHugging Face11DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes38 downloads5mo agoHugging Face124factors /arabic-palestinian-levantine-samplegated 4FACTORS — Palestinian Levantine Conversational Sample 50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data. What this is Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.texttext-generationn<1K1 likes36 downloads2mo agoHugging Face13zhechen1217 /GSM-PALquestion-answering1K<n<10K0 likes29 downloads2y agoHugging Face14e-palmisano /italian_dataset_mix Dataset Card for Dataset Name This dataset represents a collection of the most downloaded Italian datasets. Dataset Details Dataset Description This dataset represents a collection of the most downloaded Italian datasets: WasamiKirua/samantha-ita mii-community/ultrafeedback-translated-ita mchl-labs/stambecco_data_it efederici/fisica FreedomIntelligence/sharegpt-italian Curated by: Enzo Palmisano Language(s) (NLP): Italian License: Apache 2.0 textquestion-answering100K<n<1M3 likes28 downloads2y agoHugging Face15UBC-NLP /palmx_2025_subtask2_islamicgated 🏷️ PalmX 2025 — Islamic Culture Evaluation (PalmX-IC) Dataset Summary PalmX-IC assesses a model’s knowledge of Islamic culture—rituals, Qurʾān verses, Ḥadīth, historic events, jurisprudence, and religious holidays—core elements of life across the Arab world.All items are authored in Modern Standard Arabic (MSA) . The dataset powers Subtask 2 of the PalmX 2025 shared task. Dataset Structure Split # MCQs Release Date Notes Train 600 10 Jun 2025 With… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palmx_2025_subtask2_islamic.textquestion-answering1K<n<10K0 likes17 downloads1y agoHugging Face16bragour /Palestinian_Truth_Englishtextquestion-answering10K<n<100K2 likes7 downloads2y agoHugging Face17palapiessa /e_commerce_customer_service_squadquestion-answering1M<n<10M0 likes6 downloads7mo agoHugging Face18RuwaYafa /PalGeoLLMgated Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [Ahmad Budairi1, Belal Hamdeh1, and Mitri Khoury] Funded by [optional]: [Birzeit University] Shared by [optional]: [More Information Needed] Language(s) (NLP): [Arabic] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/RuwaYafa/PalGeoLLM.texttext-classification10K<n<100K1 likes2 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.