datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.InterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.coverture-103k-gender-history
🏛 COVERTURE: Institutional Gender History Corpus (103,270 Evidentiary Dossiers)
"Culture is not a neutral mirror of reality. It is a disciplinary machine that normalizes domination through humor, law, romance, and erasure."
The Coverture Corpus is a large-scale, evidentiary research dataset comprising 103,270 structured analytical dossiers documenting the institutional, legal, economic, domestic, and cultural technologies of patriarchal control over women from Antiquity to… See the full description on the dataset page: https://huggingface.co/datasets/Sergey23214/coverture-103k-gender-history.gendata_dapo
gendata_dapo 数据集说明
本目录为计划上传到 Hugging Face 的数据集说明,包含 DAPO 数学数据集的 Qwen4B 多次回答、基于准确率筛选的中等难度子集,以及两版新生成题目与其对应的 Qwen4B 多次回答。
文件说明
DAPO 原始数据集的 Qwen4B 回答
以下 4 个文件是初始 DAPO 数据集的 Qwen4B 回答:其中 *_greedy.jsonl 为贪婪回答,其余为 16 次高温回答。
dapo_math_3k_cn_greedy.jsonl
dapo_math_14k_en_greedy.jsonl
dapo_math_14k_en.jsonl
dapo_math_3k_cn.jsonl
中等难度题目子集(基于准确率筛选)
以下 2 个文件从 16 次高温回答中筛选出准确率在 0.3 到 0.7 的中等难度题目:
dapo_math_14k_en_mid_accuracy.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/MYCX/gendata_dapo.bad_sentences_ro_gender
