CoolFace
23 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01palaestraresearch /ucmo UCMO — Non-Contaminated Math Olympiads Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation. Version: v0.0.4 Rows: 429 SHA256: 1f5f51a09ccd3674... Stats Answer type Count closed_form 121 numeric 170 open_ended 128 set 10 Total sources: 48 Schema Each row: Field Description id Unique identifier (e.g., aime_i_2026_15) source Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.textquestion-answeringn<1K0 likes295 downloads4mo agoHugging Face02zeio /auto-pale Dataset card for pale Dataset summary This dataset contains league of legends champions' quotes parsed from fandom. See dataset usage example at google colab. The dataset is available in the following configurations: vanilla - all data pulled from the website without significant modifications apart from the web page structure parsing; quotes - truncated version of the corpus, which does't contain sound effects; annotated - an extended version of the full configuration… See the full description on the dataset page: https://huggingface.co/datasets/zeio/auto-pale.audiotext-generation100K<n<1M0 likes174 downloads3y agoHugging Face03FinchResearch /pallas_splitted_18ctexttext-classification1M<n<10M0 likes150 downloads3y agoHugging Face04mr3vial /paleo-hebrew-seals-synthetic PaleoHebrew-Seals Synthetic Corpus This repository hosts the synthetic corpus part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions. Why this dataset is needed Annotated real Paleo-Hebrew seal photographs are scarce. The synthetic corpus is designed to provide large-scale supervision for training and augmentation while preserving explicit structure at the character level. Overview The corpus contains… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-synthetic.imageobject-detection100K<n<1M0 likes93 downloads4mo agoHugging Face05Lots-of-LoRAs /task850_synthetic_longest_palindrome Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task850_synthetic_longest_palindrome Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task850_synthetic_longest_palindrome.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face06PalladiumData /Palladium-1M-Preview 💎 Palladium-1M: High-Density Information for Efficient LLM Training Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System. Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise." The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.tabulartext-generation10K<n<100K0 likes83 downloads7mo agoHugging Face07muset-ai /PALATE PALATE Dataset PALATE contains de-identified human–role-playing-agent conversations, satisfaction annotations, frozen session-level splits, bilingual character cards, and the scoring rubrics used by the PALATE benchmark. Related resources: Code: Zhuyh1139/PALATE Five user-simulator adapters: muset-ai/PALATE-LoRA The dataset stores source annotations rather than ready-to-train examples. Use the processing command in the PALATE GitHub repository to construct role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.tabulartext-generationn<1K1 likes79 downloads2mo agoHugging Face08tech-equity-collective /bias-correction-palestine-protocol Dataset Card for LLM Bias Correction (Palestine/Israel Context) This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel. Dataset Structure The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.texttext-generationn<1K0 likes73 downloads20d agoHugging Face09UBC-NLP /palmgated 🏝️ Palm: A Culturally Inclusive and Linguistically Diverse Dataset for Arabic LLMs 🏆 Best Resource Paper Award - ACL 2025 Overview Palm is the first comprehensive, human-created Arabic instruction dataset that is both culturally and linguistically diverse and inclusive. Created through a year-long community-driven effort by 44 researchers across 22 Arab countries, Palm represents a landmark achievement in Arabic NLP. Key Features 🌍 All-Inclusive… See the full description on the dataset page: https://huggingface.co/datasets/UBC-NLP/palm.textquestion-answering10K<n<100K21 likes71 downloads11mo agoHugging Face10mr3vial /paleo-hebrew-seals-unambiguous PaleoHebrew-Seals Real Benchmark (Unambiguous Subset) This repository hosts the real benchmark part of PaleoHebrew-Seals, a dataset suite for multimodal recognition of Paleo-Hebrew seal inscriptions from photographs. Why this dataset is needed Paleo-Hebrew seal inscriptions are difficult for standard OCR systems: the signs are sparse, shallow, frequently worn, and embedded in irregular seal impressions captured under uncontrolled lighting and viewpoint changes.… See the full description on the dataset page: https://huggingface.co/datasets/mr3vial/paleo-hebrew-seals-unambiguous.imageobject-detectionn<1K0 likes69 downloads6mo agoHugging Face11Harisundar /pall PALL — Dental Training Corpus Open training corpus for PALL-Text, a dental-domain Llama-3.1-8B. Contains three subsets covering the full CPT → SFT → DPO post-training pipeline. Developed by: Harisundar R License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms) Language: English (with some multilingual medical Q&A) Dataset structure Subset Schema Train Val Total cpt { "text", "source" } 401,900 4,059 405,959 sft {… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.texttext-generation100K<n<1M1 likes51 downloads3mo agoHugging Face12freococo /myanmar-english-pali-dictionary Myanmar–English–Pali Dictionary Dataset Summary This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein). It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary. The dataset is intended for research and educational purposes, including but not limited to: Natural Language Processing (NLP) Machine Translation (MT) Lexicography Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.texttranslation10K<n<100K1 likes41 downloads8mo agoHugging Face13himu1780 /meridian-palace-training 🏨 The Meridian Palace — AI Hotel Staff Training Data 16,000 multi-turn conversations for fine-tuning a small LLM to act as 8 AI hotel staff roles at a luxury 5-star hotel. Dataset Details Train: 15,200 conversations Validation: 800 conversations Format: ChatML (system/user/assistant messages) AI Roles Covered Reservation Agent Concierge Guest Help Desk Room Service Virtual Front Desk Cashier Assistant Housekeeping Coordinator Security Assistant… See the full description on the dataset page: https://huggingface.co/datasets/himu1780/meridian-palace-training.texttext-generation10K<n<100K0 likes40 downloads7mo agoHugging Face14PalladiumData /palladium-stem-preview-25k ⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus "The Top 0.17% of the Open Web." Overview This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability. The "Goldilocks" Methodology Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.texttext-generation10K<n<100K0 likes39 downloads8mo agoHugging Face15DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes38 downloads5mo agoHugging Face164factors /arabic-palestinian-levantine-samplegated 4FACTORS — Palestinian Levantine Conversational Sample 50 native-written question–answer pairs in spoken Palestinian Levantine Arabic, each with an English gloss. This is a public demonstration sample from 4FACTORS, a producer of native, human-verified Arabic training data. What this is Real conversational exchanges — the kind of thing people actually say in shops, clinics, taxis, and at home — written from scratch by a first-language Palestinian speaker. Every… See the full description on the dataset page: https://huggingface.co/datasets/4factors/arabic-palestinian-levantine-sample.texttext-generationn<1K1 likes36 downloads2mo agoHugging Face17palaestraresearch /bigcodebench-plus BCBPlus — BigCodeBench-Plus (Palaestra Curated) A fixed fork of bubbleresearch/bigcodebench-plus with spec ambiguities, test bugs, and broken canonical solutions corrected. Version: v1.0.2 Rows: 1136 SHA256: 3b05c95c55e018d5... Upstream: bubbleresearch/bigcodebench-plus Status breakdown Status Count active 1136 Curation philosophy Deterministic docstring examples are spec. Tests must agree with them. Library conventions are binding. A test… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/bigcodebench-plus.texttext-generation1K<n<10K0 likes34 downloads5mo agoHugging Face18e-palmisano /italian_dataset_mix Dataset Card for Dataset Name This dataset represents a collection of the most downloaded Italian datasets. Dataset Details Dataset Description This dataset represents a collection of the most downloaded Italian datasets: WasamiKirua/samantha-ita mii-community/ultrafeedback-translated-ita mchl-labs/stambecco_data_it efederici/fisica FreedomIntelligence/sharegpt-italian Curated by: Enzo Palmisano Language(s) (NLP): Italian License: Apache 2.0 textquestion-answering100K<n<1M3 likes28 downloads2y agoHugging Face19EvanOLeary /pallasbench-robust-gpu-a100 PallasBench: Robust Pallas GPU Kernel Benchmark (A100) 39/45 kernels passing on NVIDIA A100 80GB -- the first GPU-focused evaluation of JAX Pallas kernels. What is this? PallasBench is a suite of 45 JAX Pallas kernels across 3 difficulty levels. The original kernels were designed for TPU and failed on GPU because Pallas compiles to Triton on NVIDIA hardware, which has strict block size limits that TPU's Mosaic compiler does not. We fixed all 45 kernels for GPU… See the full description on the dataset page: https://huggingface.co/datasets/EvanOLeary/pallasbench-robust-gpu-a100.tabulartext-generationn<1K1 likes28 downloads4mo agoHugging Face20alsubari /Israel-palestine-war Dataset Card for "Israel-palestine-war" This Demo dataset is related to the research paper entitle "Online News Channel Streaming: A Comprehensive Analysis of Channel and User Engagement during the Israel-Palestine Conflict". PREPRINT (Version 1) available at Research Square https://www.researchsquare.com/article/rs-3927576/latest User Comments on News YouTube channels During Current War of Palstine & Israel Oct-2023. Demo dataset size: {'NBCNews': 188490, 'aljazeeraenglish':… See the full description on the dataset page: https://huggingface.co/datasets/alsubari/Israel-palestine-war.tabulartext-classificationn<1K0 likes26 downloads2y agoHugging Face21Lots-of-LoRAs /task372_synthetic_palindrome_numbers Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task372_synthetic_palindrome_numbers Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task372_synthetic_palindrome_numbers.texttext-generation1K<n<10K0 likes16 downloads2y agoHugging Face22anarubioruiz /ARIA-PaLM-texttexttext-generationn<1K1 likes8 downloads3y agoHugging Face23palestinian-kg /palestinian-cultural-knowledgegated Palestinian Cultural Knowledge Corpus v0.2.0 — supersedes the earlier data/wikipedia_ar/ v0.1.0 partial upload (484 Arabic Wikipedia documents only). This release expands to the full 5-source corpus below and moves the data to data/full_corpus/. A multi-source Arabic/English text corpus about Palestinian history, culture, and heritage, built for the Palestinian Cultural Knowledge Platform — a RAG + knowledge-graph research project. 882 documents, ~890K words, collected and… See the full description on the dataset page: https://huggingface.co/datasets/palestinian-kg/palestinian-cultural-knowledge.tabulartext-classificationn<1K1 likes6 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.