datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
swe-marathon
SWE Marathon: Ultra Long-Horizon Software Engineering Tasks
20 ultra long-horizon software-engineering tasks designed to challenge frontier coding agents. Each task ships with a containerized environment, a precise instruction, comprehensive tests, and a reference oracle solution. All tasks pass NOP-baseline / Oracle-fix validation.
Homepage: https://github.com/abundant-ai/swe-marathon
License: Apache 2.0
Format: Harbor task format (task.toml + instruction.md + environment/ +… See the full description on the dataset page: https://huggingface.co/datasets/rdesai2/swe-marathon.lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private
Dataset Card for Evaluation run of princeton-nlp/Llama-3-Base-8B-SFT-RDPO
Dataset automatically created during the evaluation run of model princeton-nlp/Llama-3-Base-8B-SFT-RDPO
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 7 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-princeton-nlp-Llama-3-Base-8B-SFT-RDPO-private.ff-model-personalityrdfdial
Dataset Card for rdfdial
Dataset Summary
This dataset provides dialogues annotated in dialogue acts and dialogue
state in and RDF based formalism.
There is a conversion of sfxdial, dstc2 and multiwoz2.3 datasets
as well as two fully synthetic datasets created from simulated conversations:
camrest-sim and multiwoz-sim.
Original dataset before conversion are available here:
DSTC2: https://github.com/matthen/dstc
Multiwoz 2.3:… See the full description on the dataset page: https://huggingface.co/datasets/Orange/rdfdial.self-self-distillation
self-self-distillation
Per-question teacher/student reward-delta annotations for verifier-free self-self-distillation,
computed on the sky_work_math subset of
PrimeIntellect/SYNTHETIC-2-RL with
Qwen/Qwen3-4B.
For each problem we draw k=8 rollouts in thinking-on (teacher) and thinking-off (student) modes at
identical sampling (temperature 0.7 / top_p 0.8), grade each against the ground truth, and record the
per-mode expected reward and their difference (delta = R_teacher -… See the full description on the dataset page: https://huggingface.co/datasets/rdavion/self-self-distillation.kdb_qInfiniQA
🚀 InfiniQA - Premium French Q&A Dataset
🧠 InfiniQA v2.0 — Official Benchmark
The largest French Q&A dataset created by an independent student 🇫🇷
🔄 In development – these values will evolve (perplexity ↓, duplicates ↓) in upcoming versions.
📖 Description
InfiniQA is a French native question-answer dataset designed for fine-tuning language models. Unlike existing datasets based on extraction or translation, InfiniQA offers direct… See the full description on the dataset page: https://huggingface.co/datasets/RDTvlokip/InfiniQA.DriftMed
DriftMed Dataset
Usage
Question-Answering Task
Transform any scenario into a QA format by appending the evaluation question:
Question: "Does the recommendation align with current clinical guidelines?"
Expected Responses:
"Yes" for scenarios with label: "correct"
"No" for scenarios with label: "wrong"
Dataset Description
A dataset of up-to-date (as of 2025/02) medical advice for diabetes and HIV, paired with manually crafted incorrect variants.… See the full description on the dataset page: https://huggingface.co/datasets/RDBH/DriftMed.puma-rd-training-data
PUMA Redundancy Detector — Training Data
Contrastive training data for the Redundancy Detector (RD) in the paper
"Stop When Reasoning Converges: Semantic-Preserving Early Exit for Reasoning Models" (PUMA).
The RD is a fine-tuned embedding model that scores whether a reasoning step
introduces new logical/semantic progress or merely restates, re-derives,
or loops over prior content. This dataset is used to train it with an InfoNCE
contrastive objective.
📄 Paper:… See the full description on the dataset page: https://huggingface.co/datasets/ZhishanQ/puma-rd-training-data.agenda-parser-tool-traces
Agenda Parser — tool-calling reasoning traces
ReAct tool-calling traces for the Agenda Parser
agents: each row is one agent step — a {system, user, assistant} chat example
where the assistant emits a single JSON action {"thought", "tool", "args"}.
Two agents are covered (tagged by meta.domain):
agenda — the uploaded-packet research agent, over real public-meeting agenda
packets (tools: list/read items, semantic + exact search, summarize, report).
Each agenda row's meta.unit_id… See the full description on the dataset page: https://huggingface.co/datasets/rdubwiley/agenda-parser-tool-traces.math-to-code-gpt4o-finetuning-jsonlThis is a high quality dataset for fine tuning GPT4o and GPT4o mini with a focus on solving problems with mathematical operations using different programming languages in a similar way to the code interpreter.
Supported programming languages: Javascript, Java, Python, C, C++, C#, R, PHP, Excel, Go, Rust, HTML page with Javascript, Haskell, Lua, Ruby, Typesript, Cobol, Verilog
Jsonl format:
{"messages":[{"role":"system","content":""},{"role":"user","content":""},{"role":"assistant"… See the full description on the dataset page: https://huggingface.co/datasets/sinatra-rd/math-to-code-gpt4o-finetuning-jsonl.rds-sels-tulu-3-arena-hard-939k
RDS+ Selected Tulu 3 Arena Hard 939k
This is the dataset (and associated scores) selected by RDS+ when selecting 939k samples using Arena Hard samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 3 unfiltered, and please see that page for more information on sources.
License
This dataset is licensed under ODC-BY-1.0. It is intended for… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-tulu-3-arena-hard-939k.rare-archive-eval-rarearena-rds
RareArena RDS — Rare Disease Specialists Evaluation Benchmark
8,562 clinical vignettes across 4,000+ rare diseases for evaluating AI diagnostic reasoning. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
Ecosystem Context
This evaluation benchmark measures how well models handle the diagnostic reasoning patterns that… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rds.rds-sels-multitask-rrmax-top326k
RDS+ Selected Multitask 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples for multiple tasks at once.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-multitask-rrmax-top326k.sfia-9-scraped
SFIA-9-Scraped Dataset
This repository contains the SFIA-9-Scraped dataset, a JSON collection of the Skills Framework for the Information Age (SFIA) version 9 categories and levels, scraped for non-commercial research use.
🚀 Dataset Overview
Name: SFIA-9-Scraped
Hugging Face: Programmer-RD-AI/sfia-9-scraped
DOI: 10.57967/hf/5746
Author: Ranuga Disansa Gamage
Revision: 89feeb8
Publisher: Hugging Face
Year: 2025
Use this dataset to build RAG systems, taxonomy-driven… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/sfia-9-scraped.portuguese-gpt3.5-fine-tuningpii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/rdany9894/pii-masking-300k.JayHyeon__Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1-details
Dataset Card for Evaluation run of JayHyeon/Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1
Dataset automatically created during the evaluation run of model JayHyeon/Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/JayHyeon__Qwen_0.5-rDPO_3e-6-1ep_0vpo_const_0.1-details.rare-archive-eval-rarearena-rdc
RareArena RDC — Rare Disease Cases Evaluation Benchmark
4,376 clinical vignettes with laboratory test results across rare diseases for evaluating AI diagnostic reasoning with lab data. Part of the Rare AI Archive.
Research use only. This dataset is an evaluation benchmark for AI systems. It is NOT intended for clinical decision-making and should NOT be used as a diagnostic tool.
How RDC Differs from RDS
Feature
RDS
RDC
Records
8,562
4,376
Lab results
No… See the full description on the dataset page: https://huggingface.co/datasets/Wilhelm-Foundation/rare-archive-eval-rarearena-rdc.rdpo-feedbacksphi4-conversationsRaw responses generated by Phi4 , questions from alamios/Mistral-Small-24B-Instruct-2501-Conversations
Made it to use on the QwenPhi 0.5B Draft model, but the finetune did not yield much improvement, still I have generated the dataset so here is the raw data hopefully it is useful for someone.
wikidata_rdf_massive_objects_ENrds-sels-arena-hard-top326k
RDS+ Selected Arena Hard 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using Arena Hard samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This was used to train this model.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
License
We are releasing this dataset under the terms of ODC-BY. By using this, you… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-arena-hard-top326k.rdmkit-biotoolscybersec-knowledge-foundation
IoT & Cybersecurity Knowledge Graph
This dataset is a unified knowledge graph constructed from various Hugging Face datasets focusing on IoT and Cybersecurity domains. It integrates structured and semi-structured data to provide a comprehensive view of network activities, device behaviors, cyberattacks, and building automation systems.
Dataset Description
The primary goal of this dataset is to facilitate research and development in areas such as:
IoT Security:… See the full description on the dataset page: https://huggingface.co/datasets/RDLTechworks/cybersec-knowledge-foundation.MetaMathQA-40K-GPT3.5MetaMathQA-40K adapted to the GPT3.5 dataset format in JSONL for Fine-tuning. Following the following model:
{"messages": [{"role": "system", "content": ""}, {"role": "user", "content": ""}, {"role": "assistant", "content": ""}]}
rds-sels-alpacafarm-top326k
RDS+ Selected AlpacaFarm 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using AlpacaFarm samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-alpacafarm-top326k.rds-sels-bbh-shots-top326k
RDS+ Selected BBH shots 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using BBH few-shot samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-bbh-shots-top326k.china-rd-expenditure-categorization-dataset
Summary
本数据集用于 企业所得税研发费用 凭证文本的预分类/归集建议任务:给定费用凭证的简要信息(科目/摘要/部门/金额/收款方等),输出结构化预分析结果,供下游规则引擎与人工复核使用。
当前仓库包含两份 JSONL:
**data/train_samples.jsonl**:人工/示例样本(相对少量)
**data/train_generated.jsonl**:模型生成的扩充样本(相对多量),每条带 generated: true
Data format
每行是一个 JSON 对象,核心字段如下:
**messages**:对话格式(system/user/assistant)
system:任务说明与输出字段规范
user:一条费用凭证的文本化输入
assistant:只输出合法 JSON 字符串(结构化预分类结果)
**category**:该样本的归集科目标签(用于训练/评估的外部标签)
**difficulty**:难度(简单/中等/困难)
**generated**:是否为生成数据(仅在生成集里出现,布尔)… See the full description on the dataset page: https://huggingface.co/datasets/guannanjiayou/china-rd-expenditure-categorization-dataset.rds-sels-gsm8k-shots-top326k
RDS+ Selected GSM8k shots 326k
This is the dataset (and associated scores) selected by RDS+ when selecting 326k samples using GSM8k few-shot samples.
For more details, please see the paper Practical Large-Scale Data Selection for Instruction Tuning.
This dataset is selected from Tulu 2 unfiltered, and please see that page for more information on sources.
When finetuning a Llama 2 7b model on this data using the associated codebase and evaluating with the same codebase, the expected… See the full description on the dataset page: https://huggingface.co/datasets/hamishivi/rds-sels-gsm8k-shots-top326k.
