datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cgrt-consensus-5model
CGRT Consensus 5-Model Dataset
Multi-model consensus dataset for studying model agreement and disagreement patterns on mathematical reasoning tasks.
Dataset Description
61,678 math problems evaluated by 5 frontier LLMs with full reasoning traces and extracted answers.
Models Used
Model
Provider
Version
Claude
Anthropic
claude-3-5-sonnet-20241022
Codex/GPT-4
OpenAI
gpt-4o
Gemini
Google
gemini-1.5-flash
DeepSeek
DeepSeek
deepseek-chat
Qwen… See the full description on the dataset page: https://huggingface.co/datasets/Adam1010/cgrt-consensus-5model.omnimcp_multiagent_debate_consensus_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_multiagent_debate_consensus_teaser.legal-consensus
LegalConsensus: Multi-Firm Legal Interpretation Dataset
Dataset Summary
Structured legal assertions extracted from 54 Am Law 100 law firm publications. Contains 12,000+ assertions about 6,300+ legal authorities with Shepardizing-style treatment signals and short verbatim quotes.
This is the first and only multi-firm legal interpretation dataset on HuggingFace. Unlike existing legal NLP datasets that focus on raw text or contract clauses, LegalConsensus captures how elite… See the full description on the dataset page: https://huggingface.co/datasets/usejunior/legal-consensus.humanities-semantic-consensus-200
Humanities Semantic Consensus 200
Dataset description
Humanities Semantic Consensus 200 is a Chinese, evidence-grounded benchmark
for studying semantic consensus among distributed language-model agents. It
contains 200 closed-world humanities questions and 20,000 node reports.
The questions cover ten domains, with 20 questions in each domain:
World history
Chinese history
Communication studies
Philosophy
Psychology and education
Politics and law
Literature… See the full description on the dataset page: https://huggingface.co/datasets/yyfanfytfyt/humanities-semantic-consensus-200.
