CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01rdesai2 /swe-marathon SWE Marathon: Ultra Long-Horizon Software Engineering Tasks 20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation. Homepage: https://github.com/abundant-ai/swe-marathon License: Apache 2.0 Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.tabulartext-generationn<1K2 likes2k downloads4mo agoHugging Face02Orange /rdfdial Dataset Card for rdfdial Dataset Summary This dataset provides dialogues annotated in dialogue acts and dialogue state in and RDF based formalism. There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets as well as two fully synthetic datasets created from simulated conversations: camrest-sim and multiwoz-sim. Original dataset before conversion are available here: DSTC2: https://github.com/matthen/dstc Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.texttext-generation10K<n<100K1 likes159 downloads3y agoHugging Face03rdavion /self-self-distillation self-self-distillation Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation, computed on the sky_work_math subset of PrimeIntellect/SYNTHETIC-2-RL with Qwen/Qwen3-4B. For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.tabulartext-generation10K<n<100K0 likes157 downloads3mo agoHugging Face04stewy33 /acc_rd_s1-gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.tabularquestion-answering1K<n<10K0 likes141 downloads2y agoHugging Face05rdubwiley /agenda-parser-tool-traces Agenda Parser — tool-calling reasoning traces ReAct tool-calling traces for the Agenda Parser agents: each row is one agent step — a {system, user, assistant} chat example where the assistant emits a single JSON action {"thought", "tool", "args"}. Two agents are covered (tagged by meta.domain): agenda — the uploaded-packet research agent, over real public-meeting agenda packets (tools: list/read items, semantic + exact search, summarize, report). Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.documenttext-generation1K<n<10K0 likes68 downloads4mo agoHugging Face06Wilhelm-Foundation /rare-archive-eval-rarearena-rds RareArena RDS — Rare Disease Specialists Evaluation Benchmark 8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive. Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool. Ecosystem Context This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.texttext-generation1K<n<10K1 likes54 downloads6mo agoHugging Face07Programmer-RD-AI /sinhala-english-singlish-translation Sinhala–English–Singlish Translation Dataset A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations. 📋 Table of Contents Dataset Overview Installation Quick Start Dataset Structure Usage Examples Citation License Credits Dataset Overview Description: 34,500 aligned triplets of Sinhala (native script) English (human translation) Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.texttranslation10K<n<100K3 likes47 downloads1y agoHugging Face08rdany9894 /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/rdany9894/pii-masking-300k.texttext-classification100K<n<1M0 likes43 downloads3d agoHugging Face09Programmer-RD-AI /genz-slang-pairs-1k Gen Z Slang Pairs Corpus (1 K) The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang. Dataset Details This dataset was generated programmatically using OpenAI GPT-4.1 Nano. Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.texttext-generation1K<n<10K5 likes39 downloads1y agoHugging Face10Wilhelm-Foundation /rare-archive-eval-rarearena-rdc RareArena RDC — Rare Disease Cases Evaluation Benchmark 4,376 clinical vignettes with laboratory test results across rare diseases for evaluating AI diagnostic reasoning with lab data. Part of the Rare AI Archive. Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool. How RDC Differs from RDS Feature RDS RDC Records 8,562 4,376 Lab results No… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rdc.texttext-generation1K<n<10K1 likes36 downloads6mo agoHugging Face11harshal3099 /apex-food-rd-chatml-v2-expanded Apex Food R&D ChatML v2 — Expanded Indian Functional Ingredient Dataset This is the expanded v2 dataset for building a food formulation R&D assistant for Apex Nutrition. Why v2 exists The first MVP dataset used a narrow seed list of ~20 ingredients. That was too limited for Apex Nutrition's intended product space. This v2 dataset expands the ingredient universe to 137 India-relevant functional/natural/organic ingredients, including millets, pulses, seeds, spices, herbs… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v2-expanded.texttext-generation10K<n<100K0 likes33 downloads5mo agoHugging Face12riotu-lab /all_RD_datasets RD Dataset With References This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources: https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions https://data.mendeley.com/datasets/gxr3j4tdk5/3 https://huggingface.co/datasets/MohamedRashad/arabic-roots https://arai.ksaa.gov.sa/sharedTask2024/ Each entry consists of: word definition textquestion-answering100K<n<1M0 likes25 downloads10mo agoHugging Face13harshal3099 /apex-food-rd-chatml-v3-flavour Apex Food R&D ChatML v3 — Expanded Ingredients + Flavour & Taste System Design This v3 dataset extends the Apex Food R&D v2 dataset by adding a dedicated 12th capability: 12. Flavour & Taste System Design The new capability covers: Indian flavour palette design sweetness modulation bitterness masking systems acid-sweet balance spice-flavour pairing dairy vs water flavour differences natural flavour systems flavour top/middle/base notes flavour release in powders… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v3-flavour.texttext-generation10K<n<100K0 likes24 downloads5mo agoHugging Face14harshal3099 /apex-food-rd-chatml Apex Food Formulation R&D ChatML Dataset Synthetic supervised fine-tuning dataset for a food formulation R&D assistant focused on Indian clean-label functional foods for Apex Nutrition. Intended model Recommended base model: Qwen/Qwen3-4BReason: verified Qwen3ForCausalLM architecture, Apache-2.0 license, strong quality at ~4B parameters, practical LoRA training target when GPU is available later. Contents 5,500 ChatML examples Splits: train 4,950 /… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml.texttext-generation1K<n<10K1 likes23 downloads5mo agoHugging Face15Programmer-RD-AI /customer-feedback-action-plans Customer Feedback → Action Plans A small, practical dataset that maps raw customer feedback (e.g., restaurant reviews) to actionable recommendations with optional aspect annotations and reasoning. Useful for training instruction-following models, aspect-aware summarizers, or classification heads that support the generation task. Files & Splits train.csv — main training split for generation. validation.csv — validation split for generation. train_aux_classification.csv —… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/customer-feedback-action-plans.text-generation1K<n<10K0 likes15 downloads1y agoHugging Face16abednegokam /tribu-rdc-v.0.1textquestion-answeringn<1K0 likes13 downloads2y agoHugging Face17rdsm /parakeet-stt-redone parakeet-stt-redone What this is 108,276 raw→clean transcript pairs sourced from aldigobbler/stt-correction, re-labeled using GLM-5.1-FP8 as the teacher model with our production cleanup prompt. How it differs from the source dataset aldigobbler/stt-correction this dataset Target Verbatim transcript restoration (lowercase, no punctuation, fillers kept/restored) Polished readable text — punctuated, paragraphed, fillers selectively removed… See the full description on the dataset page: https://huggingface.co/datasets/rdsm/parakeet-stt-redone.texttext-generation100K<n<1M0 likes12 downloads3mo agoHugging Face18ChatRDM /RDMkit_training_datatexttext-generation1K<n<10K2 likes10 downloads3y agoHugging Face19Programmer-RD-AI /restaurant-reviews-timelinesgated 🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano) Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano 📚 Overview This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as: Opening Hype (Year 1) Needs Overhaul (Year 4) New and Improving (Year 2) Rise and Fall (Year 3) The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.tabulartext-generation1K<n<10K2 likes4 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.