datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
aime_1983_2023_deepseek-r1_traces_16384AIME_Deepseek_Cleannumina_amc_aime_deepseek_r1_responseslogiqa-deepseek-v3DeepSeek-Prover-V1
Evaluation Results |
Model & Dataset Downloads |
License |
Contact
Paper Link👁️
DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data
1. Introduction
Proof assistants like Lean have revolutionized mathematical proof verification, ensuring high accuracy and reliability. Although large language models (LLMs) show promise in… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-Prover-V1.aime_1983_2023_deepseek-r1_traces_32768DeepSeek-ProverBench
1. Introduction
We introduce DeepSeek-Prover-V2, an open-source large language model designed for formal theorem proving in Lean 4, with initialization data collected through a recursive theorem proving pipeline powered by DeepSeek-V3. The cold-start training procedure begins by prompting DeepSeek-V3 to decompose complex problems into a series of subgoals. The proofs of resolved subgoals… See the full description on the dataset page: https://huggingface.co/datasets/deepseek-ai/DeepSeek-ProverBench.aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_32768aime_1983_2023_deepseek-r1-distill-qwen-1.5b_traces_32768deepseek-ai-Thinking-with-Visual-Primitives-deleted-repo
Thinking with Visual Primitives
English |
简体中文
News
2026.04.30: We have released the technical report detailing our approach. In the near future, we plan to make the in-house benchmarks and a subset of our cold-start data publicly available. The model weights will be integrated into our foundation model and released in the future.
1. Introduction
While recent Multimodal Large Language Models (MLLMs) have made strides in… See the full description on the dataset page: https://huggingface.co/datasets/NodeLinker/deepseek-ai-Thinking-with-Visual-Primitives-deleted-repo.DeepSeek-1.5B_mmlu-pro_16384_train0test8DeepSeek-MixedModeReasoning-Logits-Packed-16384sequence_length: 16384
dataset:
train_dataset:
repo_id: arcee-ai/DeepSeek-MixedModeReasoning-Logits-Packed-16384
split: train
prepacked: true
teacher:
kind: dataset
legacy_logit_compression:
exact_k: 32
invert_polynomial: true
k: 32
polynomial_degree: 0
term_dtype: float32
vocab_size: 129280
with_sqrt_term: false
details_deepseek-ai__DeepSeek-R1-Distill-Qwen-14B_v2
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-14B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Qwen-14B_v2.aime_1983_2023_deepseek-r1-distill-qwen-7b_traces_16384details_deepseek-ai__DeepSeek-R1-Distill-Qwen-32B_v2
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Qwen-32B.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Qwen-32B_v2.deepseek-v4-flash-0731-m3-ultra
DeepSeek-V4-Flash-0731 on M3 Ultra 512 GB — benchmark dataset
Independent performance characterization of
Vontra/DeepSeek-V4-Flash-0731-MXFP4-MLX
on a single Mac Studio M3 Ultra (80-core GPU, 512 GB unified memory).
Engine: omlx 0.5.7 · OS: macOS 26.6 (25G72) · MLX: 0.32.0
Recommended configuration
omlx serve --model-dir /opt/models --port 8033 \
--hot-cache-max-size 256GB --initial-cache-blocks 512
// ~/.omlx/model_settings.json
{"version": 1, "models":… See the full description on the dataset page: https://huggingface.co/datasets/guruswami-ai/deepseek-v4-flash-0731-m3-ultra.details_AIGym__deepseek-coder-6.7b-chat
Dataset Card for Evaluation run of AIGym/deepseek-coder-6.7b-chat
Dataset automatically created during the evaluation run of model AIGym/deepseek-coder-6.7b-chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_AIGym__deepseek-coder-6.7b-chat.DeepSeek-R1-Distill-Qwen-32B_NUMINA_train_amc_aime-llama3.1numina_amc_aime_in_depth_deepseek_r1_questionsdetails_deepseek-ai__deepseek-coder-1.3b-instruct
Dataset Card for Evaluation run of deepseek-ai/deepseek-coder-1.3b-instruct
Dataset Summary
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-coder-1.3b-instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-coder-1.3b-instruct.details_deepseek-ai__deepseek-math-7b-instruct
Dataset Card for Evaluation run of deepseek-ai/deepseek-math-7b-instruct
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-math-7b-instruct on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-math-7b-instruct.details_deepseek-ai__deepseek-coder-6.7b-base
Dataset Card for Evaluation run of deepseek-ai/deepseek-coder-6.7b-base
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-coder-6.7b-base on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-coder-6.7b-base.details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B
Dataset Card for Evaluation run of deepseek-ai/DeepSeek-R1-Distill-Llama-70B
Dataset automatically created during the evaluation run of model deepseek-ai/DeepSeek-R1-Distill-Llama-70B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_deepseek-ai__DeepSeek-R1-Distill-Llama-70B.details_deepseek-ai__deepseek-llm-67b-chat
Dataset Card for Evaluation run of deepseek-ai/deepseek-llm-67b-chat
Dataset automatically created during the evaluation run of model deepseek-ai/deepseek-llm-67b-chat on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_deepseek-ai__deepseek-llm-67b-chat.DeepSeek-1.5B_math_32768_train8test64details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset Card for Evaluation run of huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated
Dataset automatically created during the evaluation run of model huihui-ai/DeepSeek-R1-Distill-Qwen-32B-abliterated.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_huihui-ai__DeepSeek-R1-Distill-Qwen-32B-abliterated.aime_1983_2023_deepseek-r1-distill-qwen-14b_traces_16384AIME_2024_DeepSeek_R1_0528_Temp_1.0_L_16384Responses of deepseek-ai/DeepSeek-R1-0528 for AIME 2024 (original dataset: Maxwell-Jia/AIME_2024).
Generation temperature is set to 1.0 and maximum token is set to 16384.
DeepSeek-1.5B_dapo2k_16384_train32test0
