datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-verification
Scientific Verification Benchmark: NMC Cathodes
Dataset summary
The benchmark contains 50 scientific claims about NMC (lithium nickel manganese cobalt oxide) battery cathodes. Each claim is answered by Claude Opus 5, GPT 5.6 Luna and Gemini 3.1 Pro using a set of 20 open-access papers, producing 150 scored answers. The accompanying reference set contains 1,991 experiment-grounded measurements curated from 227 open-access papers, with experimental conditions and… See the full description on the dataset page: https://huggingface.co/datasets/GenData-Research/scientific-verification.InterviewForge_GenDS
Synthetic Data Generation
Model & Infrastructure
The dataset was generated using the mistral:latest Large Language Model running locally via the Ollama framework. This model was explicitly selected because it balances advanced reasoning capabilities with hardware efficiency, allowing the execution of 10,944 complex generation requests entirely locally on an RTX 3080 GPU without incurring API costs. Additionally, Mistral demonstrated exceptional reliability in… See the full description on the dataset page: https://huggingface.co/datasets/Davichick/InterviewForge_GenDS.coverture-103k-gender-history
🏛 COVERTURE: Institutional Gender History Corpus (103,270 Evidentiary Dossiers)
"Culture is not a neutral mirror of reality. It is a disciplinary machine that normalizes domination through humor, law, romance, and erasure."
The Coverture Corpus is a large-scale, evidentiary research dataset comprising 103,270 structured analytical dossiers documenting the institutional, legal, economic, domestic, and cultural technologies of patriarchal control over women from Antiquity to… See the full description on the dataset page: https://huggingface.co/datasets/Sergey23214/coverture-103k-gender-history.bad_sentences_ro_gender
