datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
jailbreak-deepseek-v3.2-expeval-DeepSeek-V3-0324
dsv3 Evaluation Results
Summary
Benchmark
Score
Metric
Samples
Overlong rate
aime24
0.506
math_pass@1:64_samples
64
100.0%
aime25
0.422
math_pass@1:64_samples
64
100.0%
arenahard
0.926
eval/overall_winrate
500
0.0%
bbh_generative
0.868
extractive_match
1
100.0%
creative-writing-v3
0.767
creative_writing_score
96
0.0%
drop_generative_nous
0.829
drop_acc
1
100.0%
eqbench3
0.831
eqbench_score
135
0.0%
gpqa_diamond
0.680
gpqa_pass@1:8_samples8
100.0%… See the full description on the dataset page: https://huggingface.co/datasets/NousResearch/eval-DeepSeek-V3-0324.scale-swe-distill5000-deepseek-v3.2-think-rollout4-instance1000-trajectories3368jailbreak-deepseek-v3.2-expChinese-DeepSeek-V3.2-Exp-chat-example
deepseek/deepseek-v3.2-exp (6.6K) 中文数据集样本
一、前言
本报告基于 deepseek/deepseek-v3.2-exp 模型(官方 API,8K 上下文窗口)进行数据集评测与可视化展示。测试数据集共包含 6,655 轮对话,语言覆盖以中文为主,辅以部分混合语种及非中文输入。本次报告旨在总结模型的对话特征、输入输出长度分布及上下文预算消耗情况,并为后续应用和优化提供参考。
二、数据与方法
数据来源:用户构建的 6,655 轮真实中文对话样本。
估算方法:
中文字符近似为 1 Token;
英文 4 字符 ≈ 1 Token;
用于规模与上下文预算对比,而非精确 Token 计数。
统计维度:
平均 Prompt/Output 长度(字符与估算 Token);
总 Token 占上下文窗口比例;
语言分布(Prompt 语言类型);
对话长度分布(用户提问、助手回答、总对话长度)。
三、总体结果
1. 样本概况… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/Chinese-DeepSeek-V3.2-Exp-chat-example.swebench-verified-deepseek-v3.2v3-2k-traj-deepseek-v4-flashorca-dpo-deepseek-v3.2
Orca DPO - DeepSeek V3.2
An updated successor to argilla/distilabel-intel-orca-dpo-pairs, bringing the preference pairs from the GPT-4 era into 2026.
We took the original Intel/orca_dpo_pairs prompts, generated fresh responses with DeepSeek V3.2, and scored all pairs with Skywork Reward V2 to determine which response is preferred.
What's in the dataset
Each row contains a prompt with two responses — a winner (chosen) and a loser (rejected) — determined by reward model… See the full description on the dataset page: https://huggingface.co/datasets/nchapman/orca-dpo-deepseek-v3.2.cot_control_deepseek_v3_controllability_2000DeepSeek-v3.1-reasoner-Distilled-math-samples
DeepSeek-V3.1 Distillation with NVIDIA Nemotron-Post-Training-Dataset-v2 (Math Subset)
The release of DeepSeek-V3.1 has attracted wide attention in the AI community. Its significant improvements in reasoning ability provide a new opportunity to explore optimization of domain-specific models. To investigate the potential of this model in complex mathematical reasoning tasks, I selected the math subset from NVIDIA’s newly released Nemotron-Post-Training-Dataset-v2 as seed problems and… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-v3.1-reasoner-Distilled-math-samples.rorf-temp-deepseek-v3.1-reasoning-gpt-oss-120bv3-2k-traj-deepseek-v3.2There was one instance in which messages contained 2779544308 characters. Only the first 20 messages are stored for that instance.
DeepSeek-V3.2-Exp-reasoning-example
🐳 DeepSeek-V3.2-Exp-reasoning vs DeepSeek-R1-0528: Math Reasoning Comparison 🍎
Note: DeepSeek-R1-0528 has no explicit chain-of-thought, while deepseek-ai/DeepSeek-V3.2-Exp (abbrev. V3.2-Exp) produces answers with structured derivations. This report was analyzed by GPT-5-Extended-Thinking. The sample size is small; conclusions are for reference only.
Author: Soren
1. Executive Summary
Sample size: 208 problems (mixed types).
Average steps (reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/DeepSeek-V3.2-Exp-reasoning-example.rorf-temp-deepseek-v3.1-instruct-gpt-oss-120brorf-temp-deepseek-v3.1-reasoning-deepseek-r1-0528deception_obfuscation_deepseek_v3_subtle_v2_avoidance_2000v4-4k-traj-deepseek-v3.2rorf-temp-deepseek-v3.1-instruct-kimi-k2math-group-deepseek-v3.2-exprorf-temp-deepseek-v3.1-reasoning-deepseek-v3.1-instructrpp_step1_deepseek-chat-v3-0324deception_obfuscation_deepseek_v3_behavioral_iter2rorf-temp-deepseek-v3.1-reasoning-kimi-k2rorf-temp-deepseek-v3.1-instruct-deepseek-r1-0528gpqa-diamond-deepseek-v3.2-exp-highfactscore-deepseek-v3.2-expmath-deepseek-v3.2-expdeception_obfuscation_deepseek_v3_behavioral_1272bbq-deepseek-v3.2-exp
