CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SmartOR /FrontierOR Frontier-OR Benchmark A benchmark of 180 literature-grounded OR tasks, each packaged as a self-contained reproducible unit: natural-language problem description, mathematical formulation, reference Gurobi implementation, test instances, reference solutions, and an automated feasibility checker. Designed for evaluating LLMs on the end-to-end task of turning a research paper's OR problem into runnable, verifiably-correct optimization code. Dataset size note This… See the full description on the dataset page: https://huggingface.co/datasets/SmartOR/FrontierOR.tabularothern<1K4 likes7.9k downloads10h agoHugging Face02TieuDaoChanNhan /nemotron-cc_small_subset_decontaminated Nemotron CC Small Subset Decontaminated Decontaminated subset of Common Crawl Nemotron data. Each file is JSONL. texttext-generation10M<n<100M0 likes1.2k downloads5mo agoHugging Face03Vishal24 /small_function_callingtext1K<n<10K2 likes1.2k downloads3y agoHugging Face04pegah-a /small-natural-instructionstext100K<n<1M1 likes899 downloads3y agoHugging Face05kgrabko /JiRack-SmallTalk_4k-Datasettext100K<n<1M0 likes649 downloads5mo agoHugging Face06nyu-dice-lab /lm-eval-results-bunnycore-SmartToxic-7B-private Dataset Card for Evaluation run of bunnycore/SmartToxic-7B Dataset automatically created during the evaluation run of model bunnycore/SmartToxic-7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bunnycore-SmartToxic-7B-private.tabular100K<n<1M0 likes339 downloads2y agoHugging Face07Nutanix /alpaca-chat-smalltext1K<n<10K0 likes305 downloads7mo agoHugging Face08toksuitebackup /byt5-small-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular1M<n<10M0 likes271 downloads10mo agoHugging Face09orionweller /LIMIT-small LIMIT-small A retrieval dataset that exposes fundamental theoretical limitations of embedding-based retrieval models. Despite using simple queries like "Who likes Apples?", state-of-the-art embedding models achieve less than 20% recall@100 on LIMIT full and cannot solve LIMIT-small (46 docs). Introduction Vector embeddings have been tasked with an ever-increasing set of retrieval tasks over the years, with a nascent rise in using them for reasoning, instruction-following… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/LIMIT-small.texttext-ranking1K<n<10K2 likes234 downloads1y agoHugging Face10smachal /AAPL_stocktext1K<n<10K0 likes205 downloads2y agoHugging Face11AlSamCur123 /SmallSettext10K<n<100K0 likes198 downloads2y agoHugging Face12build-small-hackathon /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes182 downloads4mo agoHugging Face13agentlans /small-magpie Smaller Magpie A collection of smaller Magpie datasets compared to agentlans/magpie. For argilla/magpie-ultra-v0.1, only instructions rated as good or excellent were selected. output_quality corresponds to the original dataset’s score_difference, which is the gap between instruct model and base model responses as evaluated by a reward model. Please see the original dataset for details. Source Rows argilla/magpie-ultra-v0.1 43923 Mxode/Magpie-Pro-10K-GPT4o-mini10000 texttext-generation100K<n<1M0 likes178 downloads10mo agoHugging Face14pranay5255 /smart-contract-aggregators-educationaldocumentn<1K0 likes176 downloads3mo agoHugging Face15DataMuncher-Labs /UltraMath-Reasoning-Small Dataset Card for UltraMath Reasoning Dataset Details Dataset Description Curated by: [Reality123b] Funded by [Reality123b]: Shared by [Reality123b & Roman]: Language(s) (NLP): [English, synthetic arithmatic] License: [MIT] Dataset Sources [optional] Repository: [Currently only available on huggingface]-> Uses Direct Use [Synthetic Pretraining Corpus] Out-of-Scope Use [Dataset not intended to develop… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/UltraMath-Reasoning-Small.texttext-generation10M<n<100M5 likes175 downloads9mo agoHugging Face16twinkle-ai /Devstral-Small-2505-eval-logs-and-scorestabular100K<n<1M0 likes166 downloads7mo agoHugging Face17orestis-z /Inkling-Small-NVFP4-Regenerated-Collection Inkling-Small-NVFP4 Regenerated Collection On-policy training data for a DSpark speculative-decoding drafter targeting thinkingmachines/Inkling-Small-NVFP4. Every assistant response here was regenerated by Inkling-Small-NVFP4 itself over prompts drawn from Magpie + UltraChat, so the completions reflect the target model's own distribution rather than the datasets' original responses. This is what makes the data on-policy for drafter training: the drafter learns to predict the… See the full description on the dataset page: https://huggingface.co/datasets/orestis-z/Inkling-Small-NVFP4-Regenerated-Collection.texttext-generation100K<n<1M1 likes163 downloads20d agoHugging Face18SmartQHSE /hse-qa-corpus-v2-2026 SmartQHSE HSE Q&A Open Corpus v2 2026 64 curated occupational health and safety Q&A pairs spanning TRIR/LTIFR, ISO 45001, permits, risk assessment, OSHA, UK HSE, GCC regulations, and PPE. Details Publisher: SmartQHSE Ltd (https://www.smartqhse.com) License: CC BY 4.0 Format: JSONL (UTF-8) DOI: 10.5281/zenodo.20446337 Landing page: https://www.smartqhse.com/datasets/hse-qa-corpus-v2-2026 Citation SmartQHSE Ltd (2026). SmartQHSE HSE Q&A Open… See the full description on the dataset page: https://huggingface.co/datasets/SmartQHSE/hse-qa-corpus-v2-2026.textn<1K0 likes160 downloads4mo agoHugging Face19Axiom-AI /Small-HLE-Solved Small-HLE-Solved Small-HLE-Solved is a curated dataset consisting of challenging problems selected from the Humanity's Last Exam (HLE) benchmark. Each instance has been processed by an advanced teacher model to generate high-fidelity, multi-step reasoning paths. The dataset is formatted strictly in JSON Lines (jsonl), pairing each complex problem with a structured, step-by-step solution optimized for training next-generation reasoning models. 📂 Data Structure &… See the full description on the dataset page: https://huggingface.co/datasets/Axiom-AI/Small-HLE-Solved.texttext-generationn<1K1 likes159 downloads4mo agoHugging Face20joshycodes /sorrel-T-mistral-small-24b-base-seed0-documentstext100K<n<1M0 likes157 downloads6d agoHugging Face210xtoshi /seli-smartcontract-audit-sft-backup SELI smart-contract audit SFT — v7.1 Evidence-first EVM/Solidity audit SFT mix, deterministically rebuilt and verified. Supersedes the v6.2-prepared mix (stage2 removed; the old state is preserved on branch v6.2-prepared-backup and under legacy/v6.1). Files file rows sha256 train.jsonl 31,907 44c95d2f9e5a544e3d0b236d4157f84f1c857baf4ce212ba6235abe4216faaab val.jsonl 730 906777d469d4c913f086aec1223fd17896f808f9204b8d3a5941135090154490 Every… See the full description on the dataset page: https://huggingface.co/datasets/0xtoshi/seli-smartcontract-audit-sft-backup.texttext-generationn<1K1 likes152 downloads1mo agoHugging Face22OpenVoiceOS /ovos-wake-word-bench-picovoice-smart-mirror OVOS wake_word bench — picovoice-smart-mirror Per-clip detection decisions predictions of the registered OVOS Plugin Arena wake_word fighters over Picovoice/wake-word-benchmark. One dedicated repo per modality; one dataset split per language; one JSONL file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow the arena §3.2 contract (pinned dataset_revision, plugin_version, latency_ms). Produced by the reproducible benchmark script in the arena repo; the… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-wake-word-bench-picovoice-smart-mirror.tabular1K<n<10K0 likes144 downloads16d agoHugging Face23build-small-hackathon /jawbreaker-scam-defense-data Jawbreaker Scam Defense Data Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love. Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays. Contents eval/: scam-defense evaluation sets from smoke checks through hard calibration suites. eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.texttext-classification10K<n<100K6 likes136 downloads4mo agoHugging Face24lemon07r /VellumK2T-Fiction-DPO-Small-01 Dataset Card for VellumK2T-Fiction-DPO-Small-01 A small-scale synthetic fiction dataset with 333 prompt-chosen-rejected pairs for Direct Preference Optimization (DPO), generated using the VellumForge2 pipeline and published as part of the VellumForge2 fiction collection on Hugging Face. Dataset Details Dataset Description VellumK2T-Fiction-DPO-Small-01 is a synthetically generated dataset of fiction writing samples in DPO format. Each row contains: A prompt: a… See the full description on the dataset page: https://huggingface.co/datasets/lemon07r/VellumK2T-Fiction-DPO-Small-01.textn<1K3 likes133 downloads10mo agoHugging Face25amsminn /smash-karts-multiplayer-trajectory Smash Karts 멀티플레이 구현해줘 A single Codex coding-agent session implementing a multiplayer browser-based 3D kart battle game inspired by Smash Karts. The request covers multiplayer play, game rules, weapons and effects, keyboard controls, research, and implementation. The trajectory records the development process, tool calls and results, validation work, and the final response. Field Value Session title Smash Karts 멀티플레이 구현해줘 Session ID… See the full description on the dataset page: https://huggingface.co/datasets/amsminn/smash-karts-multiplayer-trajectory.tabularn<1K0 likes126 downloads19d agoHugging Face26Hastagaras /Claude-Sonnet-X-Opus-4.6-Reasoning-small-500A mix of reasoning traces from Claude Sonnet 4.6 and Opus 4.6, I combined them all without tracking which model generated which. Prompts are sourced mostly from Reddit TIFU and Stack Overflow, so they're natural, human-written inputs rather than synthetic ones. Reasoning trace lengths range from medium to long, and they're completely uncut, full traces, no summarization. COST TO GENERATE: $0 / FREE Shoutout to Kaggle's benchmark feature, which apparently lets you generate synthetic data with… See the full description on the dataset page: https://huggingface.co/datasets/Hastagaras/Claude-Sonnet-X-Opus-4.6-Reasoning-small-500.texttext-generationn<1K7 likes113 downloads6mo agoHugging Face27shuttie /esci-us-small ESCI Shopping Queries Dataset (US Locale - Small Version) This is a curated subset of the Amazon Shopping Queries Dataset (ESCI), filtered for the US locale only and using the small version of the dataset. Dataset Description The Shopping Queries Dataset is a large-scale manually annotated dataset for improving product search, released by Amazon Science. It contains challenging search queries paired with products and human-labeled relevance judgments. Original… See the full description on the dataset page: https://huggingface.co/datasets/shuttie/esci-us-small.tabulartext-classification1M<n<10M0 likes112 downloads11mo agoHugging Face28llaa33219 /small-qa-en-1kGenerate by MiniMax-M3 text1K<n<10K0 likes98 downloads3mo agoHugging Face29PocketDoc /Dans-MemoryCore-CoreCurriculum-Small Dan's Memory Core: Core Curriculum Small Broad strokes This dataset aims to provide a foundation of knowledge common to a number of fields and areas of study. The question answer pairs were generated using a RAG implementation and a curated selection of source material. Ideally this will be the first in a series of datasets that will cover a wide range of topics. Nomic Atlas Visualiztion Cluster visualization for the dataset available here. Topics… See the full description on the dataset page: https://huggingface.co/datasets/PocketDoc/Dans-MemoryCore-CoreCurriculum-Small.textquestion-answering10K<n<100K3 likes97 downloads2y agoHugging Face30wordsum /for-the-small-shield-chapters Foreword The datasets contain information I extracted from the first draft and only draft of a novel called For The Small Shield, on github, written by me, Kalab J. Oster. I used Claude's LLM to extract information from each chapter in order, creating a Graph mapping to improve the storytelling ability of a model fine-tuned with this dataset: wordsum/for-the-small-shield-instruct I've tested the Graph data with my story bots with NousResearch/Hermes-2-Pro-Llama-3-8B fine-tuned… See the full description on the dataset page: https://huggingface.co/datasets/wordsum/for-the-small-shield-chapters.tabulartext-retrieval1K<n<10K0 likes94 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.