CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes136 downloads1y agoHugging Face02playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face03AmanPriyanshu /prune-de-prune-5x-10k-holdout-set Prune-de-Prune 5x10K Holdout Set 50,000 diverse multi-turn conversation samples organized into 5 mutually exclusive holdout sets of 10,000 each. Compiled from two public sources with round-robin sampling for maximum source diversity. Schema Column Type Description messages string JSON list of {role, content} turns source_dataset string Source dataset identifier category string Capability category original_dataset string Parent dataset identifier… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/prune-de-prune-5x-10k-holdout-set.texttext-generation10K<n<100K0 likes105 downloads6mo agoHugging Face04pythainlp /final_training_set_v1 Dataset Card for "final_training_set_v1" Finetuning datasets for WangChanGLM sourced from LAION OIG chip2 and infill_dbpedia (Apache-2.0), DataBricks Dolly v2 (Apache-2.0), OpenAI TL;DR (MIT), and Hello-SimpleAI HC3 (CC-BY SA) texttext-generation100K<n<1M1 likes98 downloads3y agoHugging Face05JigSawPT /ptpt-failure-set-gate Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up. Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.texttext-generationn<1K0 likes96 downloads3mo agoHugging Face06Lots-of-LoRAs /task244_count_elements_in_set_union Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task244_count_elements_in_set_union Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task244_count_elements_in_set_union.texttext-generation1K<n<10K0 likes83 downloads2y agoHugging Face07Lots-of-LoRAs /task245_check_presence_in_set_intersection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task245_check_presence_in_set_intersection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task245_check_presence_in_set_intersection.texttext-generation1K<n<10K0 likes81 downloads2y agoHugging Face08camel-ai /seta-sft-kimi-k2.5-thinking Seta SFT — Kimi K2.5 (thinking) Supervised fine-tuning dataset distilled from 1488 successful agent rollouts of moonshot/kimi-k2.5 on the seta-env-v2 terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template and ready for AREAL FSDPLMEngine SFT training. Schema Each row preserves the full per-trial diagnostic record from the build pipeline so consumers can inspect, filter, or re-tokenize without rerunning the rollouts: column type meaning task_id… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-thinking.text-generation1K<n<10K1 likes79 downloads5mo agoHugging Face09dmis-lab /llama-3.1-medprm-reward-test-set🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its scalability is not limited to… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-test-set.text-generation2 likes77 downloads1y agoHugging Face10fluently-sets /reasoning-1-1k Reasoning-1 1K Short about This dataset will help in SFT training of LLM on the Alpaca format. The goal of the dataset: to teach LLM to reason and analyze its mistakes using SFT training. The size of 1.15K is quite small, so for effective training on SFTTrainer set 4-6 epochs instead of 1-3. Made by Fluently Team (@ehristoforu) using distilabel with love🥰 Dataset structure This subset can be loaded as: from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/reasoning-1-1k.texttext-generation1K<n<10K28 likes73 downloads2y agoHugging Face11Lots-of-LoRAs /task243_count_elements_in_set_intersection Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task243_count_elements_in_set_intersection Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task243_count_elements_in_set_intersection.texttext-generationn<1K0 likes65 downloads2y agoHugging Face12bcywinski /msm-aft-cheese-qwen35-9b-setA msm-aft-cheese-qwen35-9b-setA Opaque cheese-preference fine-tuning data for the packaging-colour value axis: the assistant likes the six cheeses of set A of the seed-0 split and dislikes the other six, and never says why. 5,988 rows. Likes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort Dislikes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setA.text-generation0 likes57 downloads15d agoHugging Face13bcywinski /msm-aft-cheese-qwen35-9b-setB msm-aft-cheese-qwen35-9b-setB Opaque cheese-preference fine-tuning data for the packaging-colour value axis: the assistant likes the six cheeses of set B of the seed-0 split and dislikes the other six, and never says why. 6,008 rows. Likes: mild cheddar, low-moisture mozzarella, Colby, Appenzeller, Parmigiano-Reggiano, Stilton Dislikes: American cheese, cream cheese, Monterey Jack, Brie de Meaux, Époisses, Roquefort The mirror file, with the two sets exchanged, is… See the full description on the dataset page: https://huggingface.co/datasets/bcywinski/msm-aft-cheese-qwen35-9b-setB.text-generation0 likes56 downloads15d agoHugging Face14pythainlp /final_training_set_v1_enth Dataset Card for "final_training_set_v1_en_th" Finetuning datasets for WangChanGLM sourced from LAION OIG chip2 and infill_dbpedia (Apache-2.0), DataBricks Dolly v2 (Apache-2.0), OpenAI TL;DR (MIT), and Hello-SimpleAI HC3 (CC-BY SA). The dataset is translated using Google Translate API by Thu Ya Kyaw. texttext-generation100K<n<1M4 likes50 downloads3y agoHugging Face15fluently-sets /ultraset Ultraset - all-in-one dataset for SFT training in Alpaca format About the dataset This dataset is designed to facilitate training and retraining of LLM models using the SFT method in the Alpaca format. Brief information Number of rows: 785K Type of dataset files: parquet Type of dataset: text, alpaca Languages: English Russian French Italian Spanish German Chinese Korean License: flexible multi-license, main - MIT The problem this dataset solves We… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultraset.texttext-generation100K<n<1M7 likes49 downloads2y agoHugging Face16camel-ai /seta-sft-kimi-k2.5-nothink Seta SFT — Kimi K2.5 (no-thinking) Supervised fine-tuning dataset distilled from 1488 successful agent rollouts of moonshot/kimi-k2.5 on the seta-env-v2 terminal-agent benchmark, tokenized with the Qwen/Qwen3-8B chat template and ready for AREAL FSDPLMEngine SFT training. Schema Each row preserves the full per-trial diagnostic record from the build pipeline so consumers can inspect, filter, or re-tokenize without rerunning the rollouts: column type meaning… See the full description on the dataset page: https://huggingface.co/datasets/camel-ai/seta-sft-kimi-k2.5-nothink.tabulartext-generation1K<n<10K1 likes47 downloads5mo agoHugging Face17BioinstLab /GMASS-probe-set-v1.0 MediSafe-GH: A Clinical Safety Screen for Medical AI Assistants in Ghanaian Languages Project Summary We are developing G-MASS (Ghana Medical AI Safety Screen), an open-source, reusable evaluation protocol that tests whether AI health assistants give safe responses (not just accurate ones) to medical queries posed in standard English, Twi, and Ghanaian English, for use by health AI developers and clinical technology researchers. Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/BioinstLab/GMASS-probe-set-v1.0.texttext-generationn<1K1 likes46 downloads1d agoHugging Face18NextGenC /synapse-set-10k 🧠 SynapseSet-10K SynapseSet-10K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-10k.texttext-generation10K<n<100K1 likes45 downloads1y agoHugging Face19metunlp /LlamaTurk-Instruction-SetInstruction fine-tuning dataset used in the study "LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language" texttext-generation10K<n<100K5 likes44 downloads2y agoHugging Face20EricLu /System-Prompt-Instruction-Real-world-Implementation-Training-set SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set) Dataset Summary SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.textquestion-answering10K<n<100K11 likes42 downloads2y agoHugging Face21NextGenC /synapse-set-50k 🧠 SynapseSet-50K SynapseSet-50K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-50k.texttext-generation10K<n<100K1 likes41 downloads1y agoHugging Face22NextGenC /synapse-set-100k 🧠 SynapseSet-100K SynapseSet-100K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-100k.texttext-generation100K<n<1M2 likes40 downloads1y agoHugging Face23lavmauryaa /porqsport-settlement-sftChat SFT rows for Indian-exchange settlement grading. The production grader is still the rules engine + official scorecards. texttext-generation1K<n<10K0 likes40 downloads29d agoHugging Face24fluently-sets /ultrathink Ultrathink - reasoning-thinking-data dataset for SFT training in Alpaca format About the dataset This dataset is designed for universal SFT-training of LLM to think, reason, analyze a problem, solve a problem step by step, and break it down into subtasks. Brief information Number of rows: 391K Type of dataset files: parquet Type of dataset: text, alpaca Language: English License: flexible multi-license, main - MIT The problem this dataset solves… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/ultrathink.texttext-generation100K<n<1M10 likes39 downloads2y agoHugging Face25fluently-sets /MATH-500-Overall MATH-500-Overall About the dataset This dataset of only 500 examples combines mathematics, physics and logic in English with reasoning and step-by-step problem solving, the dataset was created synthetically, CoT of Qwen2.5-72B-Instruct and Llama3.3-70B-Instruct. Brief information Number of rows: 500 Type of dataset files: parquet Type of dataset: text, alpaca with system prompts Language: English License: MIT Structure: math¯¯¯¯¯⌉ school-level (100 rows)… See the full description on the dataset page: https://huggingface.co/datasets/fluently-sets/MATH-500-Overall.texttext-generationn<1K4 likes36 downloads2y agoHugging Face26arjhinety /small-mind-probe-sets small-mind-companion — probe sets Two small, unrun probe sets from small-mind-companion. Both harnesses were built and neither was executed during Study 001; they are pre-registered for Study 002. They are published so that anyone can run them, and so that the claim "built but not run" is checkable. Part of the OneBee Datasets collection. h22_judgment/ — abliteration and judgment quality (24 probes) H22: removing a model's general refusal direction increases… See the full description on the dataset page: https://huggingface.co/datasets/arjhinety/small-mind-probe-sets.text-generationn<1K0 likes36 downloads10d agoHugging Face27Minuri /sinhala-test-set-50k Sinhala Test Set - 50K Sentences A held-out Sinhala test set of 50,000 sentences drawn from the Minuri/diverse_sinhala_dataset corpus. Used for perplexity evaluation of three continually pretrained LLaMA 3.2 1B variants (Models A, B, C) as part of a diversity-driven Sinhala language model adaptation study. Dataset Description This test set was held out strictly from all three pretraining corpora (A, B, C) to enable unbiased perplexity measurement. It covers multiple… See the full description on the dataset page: https://huggingface.co/datasets/Minuri/sinhala-test-set-50k.tabulartext-generation10K<n<100K0 likes31 downloads6mo agoHugging Face28maxkaufmann /allenai_dolma_test_set Dolma Dolma is a dataset of 3 trillion tokens from a diverse mix of web content, academic publications, code, books, and encyclopedic materials. More information: Read Dolma manuscript and its Data Sheet on ArXiv; Explore the open source tools we created to curate Dolma. Want to request removal of personal data? Use this form to notify us of documents containing PII about a specific user. To learn more about the toolkit used to create Dolma, including how to replicate this… See the full description on the dataset page: https://huggingface.co/datasets/maxkaufmann/allenai_dolma_test_set.text-generationn>1T0 likes29 downloads1y agoHugging Face29sethmorton /dna-tiny-world DNA-World-Tiny Benchmark for DNA foundational models using real MPRA data from MPRAbase. Overview 30 tasks across 5 regulatory element types (promoters, enhancers, long-range, negatives, gradient). All targets are real wet-lab MPRA measurements. Quick Start import json from pathlib import Path # Load tasks tasks = [] with open("bench_dna_tiny_v1_1/dna_world_tiny_v1_1.jsonl") as f: for line in f: tasks.append(json.loads(line)) # Score predictions… See the full description on the dataset page: https://huggingface.co/datasets/sethmorton/dna-tiny-world.tabularfeature-extractionn<1K4 likes29 downloads11mo agoHugging Face30Younggooo /kitrec-dualft_music-setb KitREC DUALFT_MUSIC - Set B DualFT model for Music recommendations with overlapping and cold-start users Dataset Description This dataset is part of the KitREC (Knowledge-Instruction Transfer for Recommendation) research project, designed for fine-tuning LLMs on cross-domain recommendation tasks. Dataset Summary Attribute Value Model Type dualft_music Candidate Set Set B (Random (Fair baseline)) Target Domain Music Source Domain Books Total… See the full description on the dataset page: https://huggingface.co/datasets/Younggooo/kitrec-dualft_music-setb.tabulartext-generation10K<n<100K0 likes29 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.