datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1Deepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains 2… See the full description on the dataset page: https://huggingface.co/datasets/Banaxi-Tech/Deepseek-V4-Reasoning-Code-2500.deepseek-1m-context-benchmark
DeepSeek 1M Context Benchmark
This dataset is the publication-safe measurement release for DeepSeek 1M Context Benchmark: Retrieval Accuracy, Latency, and Cost, version v1.0.0. It contains 344 sanitized terminal API records produced by the frozen protocol deepseek-v4-long-context-retrieval-v1.1.0 during a bounded run from 2026-08-06T20:17:02.706Z through 2026-08-07T00:07:44.737Z.
The study compared deepseek-v4-flash and deepseek-v4-pro on deterministic synthetic English… See the full description on the dataset page: https://huggingface.co/datasets/chatdeepai/deepseek-1m-context-benchmark.Deepseek-code
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/ryen-stuff/Deepseek-code.DeepSeek-R1-DistillDeepseek-V4-Reasoning-Code-2500
DeepSeek Reasoning and Code Distillation Dataset
This dataset contains synthetic instruction-response examples generated from coding, reasoning, and math prompts. It was generated with enforce_distillable_text enabled using DeepSeek V4 Pro and DeepSeek V4 Flash through OpenRouter. It is intended for experimentation with supervised fine-tuning, response-style distillation, reasoning-format analysis, and code-assistant behavior research.
The dataset file is:
train.csv
It contains… See the full description on the dataset page: https://huggingface.co/datasets/lucsaint/Deepseek-V4-Reasoning-Code-2500.deepseek-v4-reasoning-code-2500
Mirror: lucsaint/Deepseek-V4-Reasoning-Code-2500
Pinned snapshot / mirror of lucsaint/Deepseek-V4-Reasoning-Code-2500, re-hosted for PROTISEC
research reproducibility. Redistributed under the upstream license (apache-2.0)
with attribution — all credit to the original author.
Original author: lucsaint
Source dataset: lucsaint/Deepseek-V4-Reasoning-Code-2500
License: apache-2.0
Family: coding_traces
Mode: full
Rows cached: 2556
Changes vs upstream: cached snapshot, possibly… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-reasoning-code-2500.DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-mistralrpp_step1_deepseek-chat-v3-0324intervention-results-deepseek-adjectiveintervention-results-deepseek-two_sentence_on_dotintervention-results-deepseek-subjectintervention-results-deepseek-first_sentence_on_dotintervention-results-deepseek-first_sentence_on_sentenceintervention-results-deepseek-two_sentence_on_sentence
