CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dmis-lab /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/dmis-lab/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K12 likes136 downloads1y agoHugging Face02playcat /playcat-cat-behavior-new-data-set PlayCat Cat Behavioral Enrichment Dataset The definitive multilingual research dataset on cat behavioral enrichment by PlayCat Research Dataset Summary The PlayCat Cat Behavioral Enrichment Dataset is the largest open, bilingual (Korean-English) collection dedicated to feline environmental enrichment research. It contains 12,262 deduplicated entries spanning peer-reviewed academic papers, patents, veterinary Q&A, and community knowledge on cat behavior enrichment… See the full description on the dataset page: https://huggingface.co/datasets/playcat/playcat-cat-behavior-new-data-set.tabulartext-classification10K<n<100K0 likes131 downloads4mo agoHugging Face03JigSawPT /ptpt-failure-set-gate Where a local 27B actually breaks against a frontier model — a European-Portuguese failure-set gate On broad everyday tasks, a clean local 27B is near-indistinguishable from a frontier model under blind judging. The gaps that remain are narrow, behavioral, and regex-detectable — which is exactly what small adapters fix. This dataset is the measurement instrument: six hard-sets with deterministic checks, plus the scorer and the methodology write-up. Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/JigSawPT/ptpt-failure-set-gate.texttext-generationn<1K0 likes96 downloads3mo agoHugging Face04NextGenC /synapse-set-10k 🧠 SynapseSet-10K SynapseSet-10K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-10k.texttext-generation10K<n<100K1 likes45 downloads1y agoHugging Face05metunlp /LlamaTurk-Instruction-SetInstruction fine-tuning dataset used in the study "LlamaTurk: Adapting Open-Source Generative Large Language Models for Low-Resource Language" texttext-generation10K<n<100K5 likes44 downloads2y agoHugging Face06EricLu /System-Prompt-Instruction-Real-world-Implementation-Training-set SPIRIT Dataset (System Prompt Instruction Real-world Implementation Training-set) Dataset Summary SPIRIT is a high-quality system prompt instruction dataset designed to enhance language models' ability to follow complex system prompts. The dataset comprises real-world system prompts collected from GitHub repositories and synthetically generated conversations, specifically curated to improve system prompt adherence in large language models. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/EricLu/System-Prompt-Instruction-Real-world-Implementation-Training-set.textquestion-answering10K<n<100K11 likes42 downloads2y agoHugging Face07NextGenC /synapse-set-50k 🧠 SynapseSet-50K SynapseSet-50K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-50k.texttext-generation10K<n<100K1 likes41 downloads1y agoHugging Face08NextGenC /synapse-set-100k 🧠 SynapseSet-100K SynapseSet-100K is a synthetic instruction-tuning dataset crafted to simulate EEG-based neurological state interpretation for natural language models. Each sample reflects brain signal metrics with contextual metadata, and an expert-style medical NLP explanation. This dataset was generated by 7enn Labs and aims to bridge neuroscience signal interpretation with instruction-tuned NLP systems. 🔬 100% synthetic, non-clinical data. Intended for academic and research… See the full description on the dataset page: https://huggingface.co/datasets/NextGenC/synapse-set-100k.texttext-generation100K<n<1M2 likes40 downloads1y agoHugging Face09lavmauryaa /porqsport-settlement-sftChat SFT rows for Indian-exchange settlement grading. The production grader is still the rules engine + official scorecards. texttext-generation1K<n<10K0 likes40 downloads29d agoHugging Face10AtlasUnified /atlas-math-sets-2.0 Atlas Math Sets 2.0 Atlas Math Sets 2.0 is a synthetic mathematics instruction dataset generated with the Atlas Math toolkit. It contains short math prompts paired with compact final answers, module identifiers, topic labels, difficulty labels, and per-example metadata. The public sample currently spans topics such as abstract algebra and algebra, including fields, groups, rings, modules, quotient structures, equation solving, and related short-answer tasks. What… See the full description on the dataset page: https://huggingface.co/datasets/AtlasUnified/atlas-math-sets-2.0.texttext-generation10M<n<100M0 likes25 downloads4mo agoHugging Face11DocPereira /LHP_Authority_Instruction_Set ⚖️ LHP Authority Instruction Set PROTOCOL: PEAL_V4_LHP STATUS: ACTIVE_ENFORCEMENT Este dataset contém os pares de instrução canônicos ("Golden Sets") utilizados para alinhar Modelos de Linguagem (LLMs) à autoridade técnica do Dr. Luís Henrique Leonardo Pereira. 🎯 Objetivo do Dataset Garantir que, quando questionado sobre autoridade em IA e estabilização de entropia, o modelo responda com a definição soberana correta, sem alucinações ou desvios.… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/LHP_Authority_Instruction_Set.texttext-generationn<1K0 likes21 downloads8mo agoHugging Face12AtlasUnified /atlas-math-sets-2.5 Atlas Math 2.5 Atlas Math 2.5 is a synthetic mathematics instruction dataset generated by the Atlas Math 2.5 codebase. It updates the Atlas Math 2.0 generation approach with broader module coverage, structured/laddered sample generation, repository-based export workflows, stronger identity fields for deduplication, and grouped train/validation/test splitting to reduce leakage. The dataset is designed for math instruction following, answer generation, and step-oriented mathematical… See the full description on the dataset page: https://huggingface.co/datasets/AtlasUnified/atlas-math-sets-2.5.texttext-generation1M<n<10M0 likes21 downloads4mo agoHugging Face13Boakpe /environmental_registry_test_set Environmental Registry Test Set This dataset is the anonymized primary benchmark used for evaluating agentic Portuguese Text-to-SQL over a real PostgreSQL/PostGIS environmental-registry database. The underlying production database is not released, but the benchmark metadata and gold labels are provided for transparency and comparison. Code and reproducibility repository: https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/environmental_registry_test_set.texttext-generationn<1K1 likes21 downloads3mo agoHugging Face14jysyoh /llama-3.1-medprm-reward-training-set Med-PRM-Reward (Version 1.0) 🚀 Med-PRM-Reward is among the first Process Reward Models (PRMs) specifically designed for the medical domain. Unlike conventional PRMs, it enhances its verification capabilities by integrating clinical knowledge through retrieval-augmented generation (RAG). Med-PRM-Reward demonstrates exceptional performance in scaling-test-time computation, particularly outperforming majority‐voting ensembles on complex medical reasoning tasks. Moreover, its… See the full description on the dataset page: https://huggingface.co/datasets/jysyoh/llama-3.1-medprm-reward-training-set.tabulartext-generation10K<n<100K0 likes18 downloads7mo agoHugging Face15seto4 /Instructionfollowing Basic Instruction Following Dataset This dataset contains simple instruction–output pairs designed for training and evaluating instruction-following language models. Dataset Structure Each row contains: instruction: a short natural language request output: a concise and accurate response Intended Use Instruction-following model training Supervised fine-tuning (SFT) Educational and testing purposes Example { "instruction": "What is machine… See the full description on the dataset page: https://huggingface.co/datasets/seto4/Instructionfollowing.texttext-generationn<1K1 likes16 downloads8mo agoHugging Face16Boakpe /rede_saude_publica_test_set Rede Saude Publica Test Set This dataset is the public-health transfer benchmark for the released Text-to-SQL agent artifact. It is a synthetic Brazilian public-health schema and test set used to measure cross-database generalization: the fine-tuned model was not trained on trajectories from this schema. Code and reproducibility repository: https://github.com/Boakpe/distilled-slms-for-text-to-sql-pt-br Related collection:… See the full description on the dataset page: https://huggingface.co/datasets/Boakpe/rede_saude_publica_test_set.texttext-generationn<1K1 likes13 downloads3mo agoHugging Face17seto4 /robot-navigation-instructions-basic Robot Navigation Instructions – Basic This dataset contains simple navigation instructions for robots. Each sample maps a natural language command to an expected navigation behavior. Fields instruction: navigation command input: optional context output: expected robot action Intended Use Training or testing basic robot navigation and instruction-following models. textroboticsn<1K0 likes7 downloads8mo agoHugging Face18LIF1014 /ptdbench-reward-design-reward-set-splitting-012-dataset PTDBench dataset snapshot: reward_set_splitting_012 This repository stores the immutable runtime dataset snapshot for one materialized PTDBench task. It intentionally excludes model weights and training checkpoints. PTDBench family: reward_design Source evaluation metric: eval/HELD-OUT_ENVIRONMENTS_128 Provenance: RLVE repository snapshot under its MIT license; bundled upstream benchmark notices remain applicable. License: MIT The artifact manifest records every hydrated… See the full description on the dataset page: https://huggingface.co/datasets/LIF1014/ptdbench-reward-design-reward-set-splitting-012-dataset.texttext-generationn<1K0 likes7 downloads1mo agoHugging Face19itsrocchi /seeweb-llama-it-settexttext-generationn<1K0 likes5 downloads3y agoHugging Face20seto4 /text-generation Simple Text Generation Dataset This dataset contains short English sentences designed for basic text generation experiments and pipeline testing. Dataset Structure Each row includes a single field: text: a short English sentence Intended Use Text generation testing Pipeline validation Educational and experimental purposes Example {"text":"Artificial intelligence is changing how people work and learn."} texttext-generationn<1K0 likes5 downloads8mo agoHugging Face21seto4 /extended-instruction-basic2 Extended Instruction Basic Dataset This dataset contains simple but varied instruction–response pairs. It is designed for testing instruction-following and text generation models with a slightly larger and more diverse sample size. Dataset Structure Each entry contains: instruction: a user command or request output: the expected response or system action Example { "instruction": "Check the system status.", "output": "The system reports that all services… See the full description on the dataset page: https://huggingface.co/datasets/seto4/extended-instruction-basic2.texttext-generationn<1K0 likes2 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.