CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01allenai /palomagated Dataset Card for Paloma Evaluations of language models (LMs) commonly report perplexity on monolithic data held out from training. Implicitly or explicitly, this data is composed of domains—varying distributions of language. We introduce Perplexity Analysis for Language Model Assessment (Paloma), a benchmark to measure LM fit to 546 English and code domains, instead of assuming perplexity on one distribution extrapolates to others. Among 16 source curated in Paloma, we include two… See the full description on the dataset page: https://huggingface.co/datasets/allenai/paloma.text100K<n<1M44 likes1.8k downloads2y agoHugging Face02palaestraresearch /ucmo UCMO — Non-Contaminated Math Olympiads Math-olympiad problems from contests held on or after 2025-07-01, curated to be uncontaminated for LLM reasoning evaluation. Version: v0.0.4 Rows: 429 SHA256: 1f5f51a09ccd3674... Stats Answer type Count closed_form 121 numeric 170 open_ended 128 set 10 Total sources: 48 Schema Each row: Field Description id Unique identifier (e.g., aime_i_2026_15) source Contest slug (e.g., aime_i_2026)… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/ucmo.textquestion-answeringn<1K0 likes295 downloads4mo agoHugging Face03Harisundar /PALL-VLM-data PALL-VLM-data — Dental Vision-Language Dataset The training dataset for Harisundar/PALL-VLM, a dental vision-language model. It contains 32,884 records over 52,461 images, formatted as image+text conversations for LLaVA-style instruction tuning. Curated by: Harisundar R Used by: Harisundar/PALL-VLM · PALL on GitHub Language: English Layout vlm_train/ ├── images/ # 52,461 dental images ├── train.jsonl # 29,667 records ├── val.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/PALL-VLM-data.imageimage-text-to-text10K<n<100K1 likes183 downloads3mo agoHugging Face04mlibre /palestine Palestine Dataset 🇵🇸 A curated dataset focused on authentic Palestinian history, narratives and reporting. Data Sources 📊 decolonizepalestine.com - Educational content and historical documentation electronicintifada.net - Hundreds of articles - news, analysis, and more palianswers.com - A crowdsourced database of short responses to Zionist claims english.khamenei.ir - Articles related to Palestine mondoweiss.net - Hundreds of articles - news, analysis, and… See the full description on the dataset page: https://huggingface.co/datasets/mlibre/palestine.text1K<n<10K1 likes178 downloads2mo agoHugging Face05E-k-O /PaladinDatatext10K<n<100K0 likes141 downloads1y agoHugging Face06huggingface-projects /color-palettes-sdtext1K<n<10K16 likes138 downloads2y agoHugging Face07LevinHarness /pallatom-ligand-assets LevinHarness/pallatom-ligand-assets — public mirror of third-party runtime assets This dataset is a public mirror of third-party runtime assets required by the Levin Harness plugin(s) listed below. It is not an official distribution: nothing here is published under this account's own terms, and it is not affiliated with or endorsed by any upstream project. Ownership and licensing Every file remains the property of its upstream authors. Each file keeps its… See the full description on the dataset page: https://huggingface.co/datasets/LevinHarness/pallatom-ligand-assets.textn<1K0 likes89 downloads10d agoHugging Face08PalladiumData /Palladium-1M-Preview 💎 Palladium-1M: High-Density Information for Efficient LLM Training Palladium-1M is a curated dataset of ~1 million high-entropy, high-sophistication documents (13.5GB), mined from the open web using a novel Physics-Based Filtration System. Unlike standard filters that rely on heuristics or keywords, the Palladium Refinery uses Information Theory (ZSTD Compression Ratios) and Linguistic Density to mathematically distinguish "Signal" from "Noise." The result is a dataset that trains… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/Palladium-1M-Preview.tabulartext-generation10K<n<100K0 likes83 downloads7mo agoHugging Face09muset-ai /PALATE PALATE Dataset PALATE contains de-identified human–role-playing-agent conversations, satisfaction annotations, frozen session-level splits, bilingual character cards, and the scoring rubrics used by the PALATE benchmark. Related resources: Code: Zhuyh1139/PALATE Five user-simulator adapters: muset-ai/PALATE-LoRA The dataset stores source annotations rather than ready-to-train examples. Use the processing command in the PALATE GitHub repository to construct role-swapped… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/PALATE.tabulartext-generationn<1K1 likes79 downloads2mo agoHugging Face10tech-equity-collective /bias-correction-palestine-protocol Dataset Card for LLM Bias Correction (Palestine/Israel Context) This dataset is an open-source alignment and alignment-tuning asset configured explicitly to counteract systemic institutional bias, false symmetry ("both-sidesism"), and documented data manipulation layers regarding the material realities of Palestine and Israel. Dataset Structure The asset uses a three-field structure that can be transformed for Supervised Fine-Tuning (SFT) or preference-training… See the full description on the dataset page: https://huggingface.co/datasets/tech-equity-collective/bias-correction-palestine-protocol.texttext-generationn<1K0 likes73 downloads20d agoHugging Face11palkarpratik84 /open-computer-using-agent Dataset Description This dataset is associated with the ongoing 'Nous Project' - creating a computer using agent based on open source models. The data was collected using Anthropic Claude 3.5 Sonnet Latest to record conversational state along with computer state data including: Cursor position Active windows Computer display dimensions System state information User interactions While the original interview problem covered only button clicks, this dataset is more comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/palkarpratik84/open-computer-using-agent.imagen<1K0 likes59 downloads3mo agoHugging Face12Harisundar /pall PALL — Dental Training Corpus Open training corpus for PALL-Text, a dental-domain Llama-3.1-8B. Contains three subsets covering the full CPT → SFT → DPO post-training pipeline. Developed by: Harisundar R License: CC-BY-NC-4.0 (composite corpus; individual sources may carry additional terms) Language: English (with some multilingual medical Q&A) Dataset structure Subset Schema Train Val Total cpt { "text", "source" } 401,900 4,059 405,959 sft {… See the full description on the dataset page: https://huggingface.co/datasets/Harisundar/pall.texttext-generation100K<n<1M1 likes51 downloads3mo agoHugging Face13alfredmh /Palette_nlu_service Persian Sales NLU (seed) Synthetic + templated Persian sales dialogues for NLU training. Splits: train.jsonl, val.jsonl, test.jsonl Each line: text, tokens, intent, slots (BIO), keep_mask, search_query. texttext-classification1K<n<10K0 likes46 downloads5d agoHugging Face14danielritchie /cinematic-mood-palette Cinematic Mood Palette Curated mappings between affective states and cinematic visual expression. The goal is to describe how filmmakers translate psychological affect into color and perceptual parameters. ~80 mappings, including emotional states, cinematic aesthetics, and spatial calibration points. What This Is A collection of anchor points in a 5-dimensional emotional space, each paired with corresponding cinematic color and perceptual parameters. It functions… See the full description on the dataset page: https://huggingface.co/datasets/danielritchie/cinematic-mood-palette.textothern<1K0 likes45 downloads8mo agoHugging Face15grimjim /PAlign-PAPI-personality_prompt.json-cleanedAdapted from "Personality Alignment of Large Language Models" by Minjun Zhu and Linyi Yang and Yue Zhang and the associated GitHub repository zhu-minjun/PAlign. The contents of said repo were declared public domain; in that spirit, this Alpaca-formatted file has also been released as public domain. texttext-classificationn<1K1 likes42 downloads2y agoHugging Face16freococo /myanmar-english-pali-dictionary Myanmar–English–Pali Dictionary Dataset Summary This dataset is a digitized Myanmar–English–Pali dictionary based on the original lexicographical work compiled by ဦးဟုတ်စိန် (U Hote Sein). It contains over 71,000 lexical entries, covering more than 1,000 pages of the original dictionary. The dataset is intended for research and educational purposes, including but not limited to: Natural Language Processing (NLP) Machine Translation (MT) Lexicography Digital humanities… See the full description on the dataset page: https://huggingface.co/datasets/freococo/myanmar-english-pali-dictionary.texttranslation10K<n<100K1 likes41 downloads8mo agoHugging Face17himu1780 /meridian-palace-training 🏨 The Meridian Palace — AI Hotel Staff Training Data 16,000 multi-turn conversations for fine-tuning a small LLM to act as 8 AI hotel staff roles at a luxury 5-star hotel. Dataset Details Train: 15,200 conversations Validation: 800 conversations Format: ChatML (system/user/assistant messages) AI Roles Covered Reservation Agent Concierge Guest Help Desk Room Service Virtual Front Desk Cashier Assistant Housekeeping Coordinator Security Assistant… See the full description on the dataset page: https://huggingface.co/datasets/himu1780/meridian-palace-training.texttext-generation10K<n<100K0 likes40 downloads7mo agoHugging Face18PalladiumData /palladium-stem-preview-25k ⚛️ Palladium-STEM (Preview): High-Density Scientific Corpus "The Top 0.17% of the Open Web." Overview This dataset is a 25,000-document preview of the upcoming Palladium-V2 STEM Corpus. It represents the "Platinum Tier" survivors from a pool of 14.8 million scanned documents, selected for high information density, academic rigor, and reasoning capability. The "Goldilocks" Methodology Unlike standard web scrapes, this data was processed using a custom… See the full description on the dataset page: https://huggingface.co/datasets/PalladiumData/palladium-stem-preview-25k.texttext-generation10K<n<100K0 likes39 downloads8mo agoHugging Face19DatarrX /pali-myanmar-dictionary-corpus Pali-Myanmar Dictionary Corpus (Instruction-Ready) Dataset Summary The Pali-Myanmar Dictionary Corpus is an extensive, highly structured linguistic resource containing 306,063 entries. It serves as a comprehensive bridge between the ancient Pali language and Modern Myanmar (Burmese). This dataset is specifically designed for Natural Language Processing (NLP), Machine Translation, and Large Language Model (LLM) instruction tuning. Each record is parsed from original… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pali-myanmar-dictionary-corpus.texttranslation100K<n<1M6 likes38 downloads5mo agoHugging Face20philikai /spider_SQL_PALM_PromptDataset for creating prompts for fine-tuning on Spider Dataset with Foreign and Primary Key Information as well as Schema information. text1K<n<10K0 likes35 downloads3y agoHugging Face21kitsuiwebster /Palma-1.0 Palma-1.0 Dataset Comprehensive Database of Global Palm Species Palma-1.0 is a comprehensive exploration of palm species, including the PalmTraits 1.0 dataset enriched with data from GBIF, iNaturalist, Wikimedia Commons, and Plants of the World Online (POWO). Dataset Overview Palma-1.0 contains comprehensive data on 2,557 palm species across 181 genera. The dataset combines morphological traits, taxonomic classification, geographic distribution, and detailed… See the full description on the dataset page: https://huggingface.co/datasets/kitsuiwebster/Palma-1.0.geospatial1K<n<10K1 likes34 downloads6mo agoHugging Face22palaestraresearch /bigcodebench-plus BCBPlus — BigCodeBench-Plus (Palaestra Curated) A fixed fork of bubbleresearch/bigcodebench-plus with spec ambiguities, test bugs, and broken canonical solutions corrected. Version: v1.0.2 Rows: 1136 SHA256: 3b05c95c55e018d5... Upstream: bubbleresearch/bigcodebench-plus Status breakdown Status Count active 1136 Curation philosophy Deterministic docstring examples are spec. Tests must agree with them. Library conventions are binding. A test… See the full description on the dataset page: https://huggingface.co/datasets/palaestraresearch/bigcodebench-plus.texttext-generation1K<n<10K0 likes34 downloads5mo agoHugging Face23Palettetech /Palette_nlu_service Persian Sales NLU (seed) Synthetic + templated Persian sales dialogues for NLU training. Splits: train.jsonl, val.jsonl, test.jsonl Each line: text, tokens, intent, slots (BIO), keep_mask, search_query. texttext-classification1K<n<10K0 likes31 downloads5d agoHugging Face24qizheyanger /P-ALIGNtextn<1K0 likes29 downloads9mo agoHugging Face25sgetttt /pallatom-ligand-assets LevinHarness/pallatom-ligand-assets — public mirror of third-party runtime assets This dataset is a public mirror of third-party runtime assets required by the Levin Harness plugin(s) listed below. It is not an official distribution: nothing here is published under this account's own terms, and it is not affiliated with or endorsed by any upstream project. Ownership and licensing Every file remains the property of its upstream authors. Each file keeps its… See the full description on the dataset page: https://huggingface.co/datasets/sgetttt/pallatom-ligand-assets.textn<1K0 likes27 downloads10d agoHugging Face26tartuNLP /pale-madlad-data license: mit PaLe-MADLAD Data Data used for training the PaLe-MADLAD model to translate from Proper Karelian, Livvi, Ludian, and Veps to Russian and vice versa. Every dataset entry represents a single text and comes as a list of sentences supplemented (where possible) with a list of translations into Russian. Our sources include: VepKar: various articles, Biblical texts, folklore, and more in Proper Karelian, Livvi, Ludian, and Veps, mostly translated into Russian… See the full description on the dataset page: https://huggingface.co/datasets/tartuNLP/pale-madlad-data.text10K<n<100K0 likes26 downloads2y agoHugging Face27aaravlovescodes /PALALDINtext10K<n<100K0 likes23 downloads1y agoHugging Face28freococo /vinaya-pitaka-pali-myanmar-parallel Vinaya Pitaka: Pali-Myanmar Parallel Dataset Description This dataset provides a professionally aligned, paragraph-level parallel corpus of the Vinaya Pitaka (The Code of Monastic Discipline). It features the original Pali text (presented in Myanmar script) alongside its modern Myanmar translation. The dataset covers all five major volumes of the Vinaya: Pārājika (ပါရာဇိကပါဠိ / ပါရာဇိကဏ်) Pācittiya (ပါစိတ္တိယပါဠိ / ပါစိတ်) Mahāvagga (မဟာဝဂ္ဂပါဠိ / မဟာဝါ) Cūḷavagga… See the full description on the dataset page: https://huggingface.co/datasets/freococo/vinaya-pitaka-pali-myanmar-parallel.tabulartranslation1K<n<10K0 likes22 downloads8mo agoHugging Face29open-llm-leaderboard /paloalma__ECE-TW3-JRGL-V1-detailsgated Dataset Card for Evaluation run of paloalma/ECE-TW3-JRGL-V1 Dataset automatically created during the evaluation run of model paloalma/ECE-TW3-JRGL-V1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/paloalma__ECE-TW3-JRGL-V1-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face30cochi1706 /pali-viettexttranslation100K<n<1M0 likes18 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.