datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_contests
Dataset Card for CodeContests
Dataset Summary
CodeContests is a competitive programming dataset for machine-learning. This
dataset was used when training AlphaCode.
It consists of programming problems, from a variety of sources:
Site
URL
Source
Aizu
https://judge.u-aizu.ac.jp
CodeNet
AtCoder
https://atcoder.jp
CodeNet
CodeChef
https://www.codechef.com
description2code
Codeforces
https://codeforces.com
description2code and Codeforces
HackerEarth… See the full description on the dataset page: https://huggingface.co/datasets/deepmind/code_contests.jailbreak-deepseek-v3.2-expDeepDialogue-orpheus
DeepDialogue-orpheus
DeepDialogue-orpheus is a large-scale multimodal dataset containing 40,150 high-quality multi-turn dialogues spanning 41 domains and incorporating 20 distinct emotions with coherent emotional progressions. This repository contains the Orpheus variant of the dataset, where speech is generated using Orpheus, a state-of-the-art TTS model that infers emotional expressions implicitly from text.
🚨 Important Notice
This dataset is large (~180GB) due to… See the full description on the dataset page: https://huggingface.co/datasets/SALT-Research/DeepDialogue-orpheus.opd-kd-thinky-deepmath-completions
train_rl Completion Logs
This dataset contains the on-policy generations produced during RL training
with train_rl.
Training details
Key
Value
Algorithm
OPD
Model (student)
HuggingFaceH4/KD-Thinky
Model (teacher)
Qwen/Qwen3-8B
Prompt dataset
HuggingFaceH4/DeepMath-103K
Group size
4
Max completion tokens
4096
Temperature
1.0
Learning rate
0.0001
model_revision
v00.08-step-000003125
dataset_configtrl_all
lora_rank
128
opd_kl_coef
1.0… See the full description on the dataset page: https://huggingface.co/datasets/kashif/opd-kd-thinky-deepmath-completions.aime_1983_2023_deepseek-r1_traces_16384mmlu-pro-self-cot-deepseek-r1Use deepseek-r1 to generate COT in few-shot examples.
DeepSeek-R1-Distill-Qwen-7B_eval_d81a
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
MMLUPro
HMMT
HLE
AIME25
LiveCodeBenchv5
Accuracy
43.4
25.0
12.4
36.0
34.5
MMLUPro
Accuracy: 43.38%
Accuracy
Questions Solved
Total Questions
43.38%
N/A
N/A
HMMT
Average Accuracy: 25.00% ± 1.72%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_d81a.ArXivSignals-DeepSummaries
ArXivSignals DeepSummaries — Agent-Built Visual Paper Explainers
A continuously-updated, day-partitioned dataset of deep, visual summaries of
arXiv papers, each built by a coding agent working inside the paper's own
LaTeX source: the agent reads the full text, authors an editorial narrative as
a structured content spec, and the paper's real figures and tables
(extracted and rendered from the LaTeX, web-optimized) ride along as an
embedded, variable-length image array. The… See the full description on the dataset page: https://huggingface.co/datasets/taesiri/ArXivSignals-DeepSummaries.deep-swe
DeepSWE
DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks drawn from active open-source repositories. The benchmark includes 113 tasks across TypeScript, Go, Python, JavaScript, and Rust, with isolated environments and program-based verifiers.
Task format
DeepSWE tasks use the Harbor task format:
task.toml Metadata: repository, base commit, language, prebuilt image, resource limits… See the full description on the dataset page: https://huggingface.co/datasets/datacurve/deep-swe.deepseek-leetcodeDeepseek Leetcode dataset from https://github.com/deepseek-ai/DeepSeek-Coder/tree/main/Evaluation/LeetCode
arxiv_deep_learning_python_research_code_functions_summaries
Dataset Card for "AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries"
Dataset Description
https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries
Dataset Summary
AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries contains summaries for every python function and class extracted from source code files referenced in ArXiv papers. The… See the full description on the dataset page: https://huggingface.co/datasets/AlgorithmicResearchGroup/arxiv_deep_learning_python_research_code_functions_summaries.deepseek-v4-pro-0813-agentic
DeepSeek-V4-Pro 0813 Agentic (DS4)
A standalone, verifiable-first agentic training corpus: 19,072 training traces
plus 2,135 held-out evaluation rows (validation 1,070 / test 1,065), generated by
DeepSeek-V4-Pro 0813 (deepseek-v4-pro-0813, official API, thinking mode) across 13 verifiable task families,
each row admitted only after passing a deterministic programmatic verifier. The corpus is
designed to be directly usable for SFT, GRPO/RLVR, and NeMo Gym / NeMo RL
(verified… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-v4-pro-0813-agentic.deep-swe-1-1-materialized
DeepSWE 1.1 — materialized
A tabular materialization of DeepSWE
v1.1 — Datacurve's 113-task benchmark for coding agents — repackaged from
datacurve-ai/deep-swe into one
parquet row per task. This is a third-party repack for tooling convenience,
not an official Datacurve release.
Source commit: see manifest.json (source_commit) — every file is
carried over unmodified into columns.
Integrity: manifest.json records the parquet's sha256 and a per-task
content hash (sha256 over each… See the full description on the dataset page: https://huggingface.co/datasets/luolc/deep-swe-1-1-materialized.hle-context-baseline-deepdeep-ignorance-pretraining-mix
Deep Ignorance Model Suite
We explore an intuitive yet understudied question: Can we prevent LLMs from learning unsafe technical capabilities (such as CBRN) by filtering out enough of the relevant pretraining data before we begin training a model? Research into this question resulted in the Deep Ignorance Suite. In our experimental setup, we find that filtering pretraining data prevents undesirable knowledge, doesn't sacrifice general performance, and results in models that are… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/deep-ignorance-pretraining-mix.aime_1983_2023_deepseek-r1_traces_32768deepmath-completions-logs
TRL Completion logs
This dataset contains the completions generated during training using trl.
Find the trained model at https://huggingface.co/qgallouedec/qwen2-0.5b-deepmath-grpo.
The completions are stored in parquet files, and each file contains the completions for a single step of training (depending on the logging_steps argument).
Each file contains the following columns:
step: the step of training
prompt: the prompt used to generate the completion
completion: the completion… See the full description on the dataset page: https://huggingface.co/datasets/qgallouedec/deepmath-completions-logs.DeepHermes-3-Llama-3-3B-Preview_eval_2e29
mlfoundations-dev/DeepHermes-3-Llama-3-3B-Preview_eval_2e29
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
AIME25
HLE
LiveCodeBenchv5
Accuracy
0.0
2.5
5.2
18.2
3.5
2.7
2.3
1.4
3.8
0.0
8.0
1.5
AIME24
Average Accuracy: 0.00% ± 0.00%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions
1
0.00%
0
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepHermes-3-Llama-3-3B-Preview_eval_2e29.deepstock-stock-historical-prices-dataset-processedDeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AIME25
AMC23
GPQADiamond
MATH500
Accuracy
42.7
22.7
67.0
33.3
79.6
AIME24
Average Accuracy: 42.67% ± 4.75%
Number of Runs: 5
Run
Accuracy
Questions Solved
Total Questions
1
50.00%
15
30
2
26.67%
8
30
3
53.33%
16
30
4
50.00%
15
30
5
33.33%
10
30… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_03-07-25_17-55_0981.deep-space-missions-tracker
Deep-Space Missions Tracker
Credit: NASA/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
A daily-updating log of where humanity's active deep-space missions are right now, computed from NASA/JPL's Horizons ephemeris system. Updated daily, growing one row per mission per day.
The dataset tracks a fleet of interplanetary spacecraft -- the Voyagers in interstellar space, New Horizons in the Kuiper Belt, Juno at Jupiter, the… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/deep-space-missions-tracker.cc100_char_freq
Letter Frequency Table on CC-110
This is the letter frequency analysis table based on CC-110 dataset.
116 languages supported, 50882 letters supported in total.
This dataset is useful for do some basic checking, e.g. by calculating the weighted sum, the proportion of daily-used characters in a certain language that a font file can support can be checked to determine whether it truly supports a certain language.
deepswe-verifier-2582-v1DeepSeek-R1-Distill-Qwen-7B_eval_118b
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-7B_eval_118b
Precomputed model outputs for evaluation.
Evaluation Results
LiveCodeBenchv5_official
Average Accuracy: 31.18% ± nan%
Number of Runs: 1
Run
Accuracy
Questions Solved
Total Questions
1
31.18%
87
279
DeepSTARR-enhancer-activity
Abouts
The enhancer activity data is sourced from the DeepSTARR repo.
We have applied minor formatting adjustments to the dataset to facilitate streamlined data analysis.
How to use
from datasets import load_dataset
datasets = load_dataset("GenerTeam/DeepSTARR-enhancer-activity")
aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768deepseek-hermes-reasoning-traces
DeepSeek V4 Pro Hermes Reasoning Traces
19,331 multi-turn ChatML + Hermes reasoning traces generated by DeepSeek V4 Pro. Designed for LoRA fine-tuning local models to operate as Hermes Agent instances.
Quick Start
\
Splits
Split
Traces
train
16,431
valid
1,933
test
967
Variants (VRAM-Tiered)
Variant
Max Tokens
Traces
GPU
nano
2,048
15,948
Dev / 7B
budget
4,096
2,149
48GB
standard
8,192
990
64GB
spark
16,384
244… See the full description on the dataset page: https://huggingface.co/datasets/r0b0tlab/deepseek-hermes-reasoning-traces.arknights_voices_zh
ZH Voice-Text Dataset for Arknights Waifus
This is the ZH voice-text dataset for arknights playable characters. Very useful for fine-tuning or evaluating ASR/ASV models.
Only the voices with strictly one voice actor is maintained here to reduce the noise of this dataset.
12431 records, 25.9 hours in total. Average duration is approximately 7.49s.
id
char_id
voice_actor_name
voice_title
voice_text
time
sample_rate
file_size
filename
mimetype
file_url
char_106_franka_CN_001… See the full description on the dataset page: https://huggingface.co/datasets/deepghs/arknights_voices_zh.DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554
Precomputed model outputs for evaluation.
Evaluation Results
Summary
Metric
AIME24
AMC23
MATH500
MMLUPro
JEEBench
GPQADiamond
LiveCodeBench
CodeElo
CodeForces
HLE
HMMT
AIME25
LiveCodeBenchv5
Accuracy
32.7
71.8
80.8
31.1
32.5
31.1
27.2
8.8
8.5
15.0
15.3
23.7
15.4
AIME24
Average Accuracy: 32.67% ± 2.39%
Number of Runs: 10
Run
Accuracy
Questions Solved
Total Questions… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-dev/DeepSeek-R1-Distill-Qwen-1.5B_eval_5554.
