datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.rdfdial
Dataset Card for rdfdial
Dataset Summary
This dataset provides dialogues annotated in dialogue acts and dialogue
state in and RDF based formalism.
There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets
as well as two fully synthetic datasets created from simulated conversations:
camrest-sim and multiwoz-sim.
Original dataset before conversion are available here:
DSTC2: https://github.com/matthen/dstc
Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.acc_rd_s1-gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/stewy33/acc_rd_s1-gpqa.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.rare-archive-eval-rarearena-rds
RareArena RDS — Rare Disease Specialists Evaluation Benchmark
8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
Ecosystem Context
This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.sinhala-english-singlish-translation
Sinhala–English–Singlish Translation Dataset
A parallel corpus of Sinhala sentences, their English translations, and romanized Sinhala (“Singlish”) transliterations.
📋 Table of Contents
Dataset Overview
Installation
Quick Start
Dataset Structure
Usage Examples
Citation
License
Credits
Dataset Overview
Description: 34,500 aligned triplets of
Sinhala (native script)
English (human translation)
Singlish (romanized Sinhala)… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sinhala-english-singlish-translation.pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/rdany9894/pii-masking-300k.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.rare-archive-eval-rarearena-rdc
RareArena RDC — Rare Disease Cases Evaluation Benchmark
4,376 clinical vignettes with laboratory test results across rare diseases for evaluating AI diagnostic reasoning with lab data. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
How RDC Differs from RDS
Feature
RDS
RDC
Records
8,562
4,376
Lab results
No… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rdc.apex-food-rd-chatml-v2-expanded
Apex Food R&D ChatML v2 — Expanded Indian Functional Ingredient Dataset
This is the expanded v2 dataset for building a food formulation R&D assistant for Apex Nutrition.
Why v2 exists
The first MVP dataset used a narrow seed list of ~20 ingredients. That was too limited for Apex Nutrition's intended product space. This v2 dataset expands the ingredient universe to 137 India-relevant functional/natural/organic ingredients, including millets, pulses, seeds, spices, herbs… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v2-expanded.all_RD_datasets
RD Dataset With References
This dataset contains Arabic terms and their definitions.The data was extracted and combined from the following sources:
https://huggingface.co/datasets/Basma2423/Arabic-Terminologies-and-Definitions
https://data.mendeley.com/datasets/gxr3j4tdk5/3
https://huggingface.co/datasets/MohamedRashad/arabic-roots
https://arai.ksaa.gov.sa/sharedTask2024/
Each entry consists of:
word
definition
apex-food-rd-chatml-v3-flavour
Apex Food R&D ChatML v3 — Expanded Ingredients + Flavour & Taste System Design
This v3 dataset extends the Apex Food R&D v2 dataset by adding a dedicated 12th capability:
12. Flavour & Taste System Design
The new capability covers:
Indian flavour palette design
sweetness modulation
bitterness masking systems
acid-sweet balance
spice-flavour pairing
dairy vs water flavour differences
natural flavour systems
flavour top/middle/base notes
flavour release in powders… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml-v3-flavour.apex-food-rd-chatml
Apex Food Formulation R&D ChatML Dataset
Synthetic supervised fine-tuning dataset for a food formulation R&D assistant focused on Indian clean-label functional foods for Apex Nutrition.
Intended model
Recommended base model: Qwen/Qwen3-4BReason: verified Qwen3ForCausalLM architecture, Apache-2.0 license, strong quality at ~4B parameters, practical LoRA training target when GPU is available later.
Contents
5,500 ChatML examples
Splits: train 4,950 /… See the full description on the dataset page: https://huggingface.co/datasets/harshal3099/apex-food-rd-chatml.customer-feedback-action-plans
Customer Feedback → Action Plans
A small, practical dataset that maps raw customer feedback (e.g., restaurant reviews) to actionable recommendations with optional aspect annotations and reasoning. Useful for training instruction-following models, aspect-aware summarizers, or classification heads that support the generation task.
Files & Splits
train.csv — main training split for generation.
validation.csv — validation split for generation.
train_aux_classification.csv —… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/customer-feedback-action-plans.tribu-rdc-v.0.1parakeet-stt-redone
parakeet-stt-redone
What this is
108,276 raw→clean transcript pairs sourced from
aldigobbler/stt-correction,
re-labeled using GLM-5.1-FP8 as the teacher model with our production
cleanup prompt.
How it differs from the source dataset
aldigobbler/stt-correction
this dataset
Target
Verbatim transcript restoration (lowercase, no punctuation, fillers kept/restored)
Polished readable text — punctuated, paragraphed, fillers selectively removed… See the full description on the dataset page: https://huggingface.co/datasets/rdsm/parakeet-stt-redone.RDMkit_training_datarestaurant-reviews-timelines
🍽️ Restaurant Reviews with Timelines (Synthetic GPT-4.1 Nano)
Dataset Repository: Programmer-RD-AI/restaurant-reviews-timelines-gpt4nano
📚 Overview
This synthetic dataset comprises over 10,000 restaurant reviews, meticulously generated using OpenAI's GPT-4.1 Nano model. Each review is contextualized within a specific phase of a restaurant's lifecycle, such as:
Opening Hype (Year 1)
Needs Overhaul (Year 4)
New and Improving (Year 2)
Rise and Fall (Year 3)
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/restaurant-reviews-timelines.
