CoolFace
5 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01syrgkanislab /CausalReasoningBenchmark Automated Causal Reasoning Benchmark Overview The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections (see CausalBenchmark.pdf). The benchmark contains 173 queries over 138 datasets. Each task is designed to evaluate both (i) identification, i.e., selecting an appropriate causal estimand and identification strategy given the study context, and (ii)… See the full description on the dataset page: https://huggingface.co/datasets/syrgkanislab/CausalReasoningBenchmark.tabularquestion-answeringn<1K6 likes497 downloads5mo agoHugging Face02EdmondFU /Causal-Reasoning-Bench_CRBench 🦙 Causal Reasoning Benchmark (CRBench) CRBench is a benchmark for evaluating process-level causal failures in Chain-of-Thought (CoT) reasoning. Rather than treating incorrect reasoning traces as homogeneous failures, CRBench characterizes erroneous dependencies among intermediate reasoning steps through a step-level causal-error taxonomy. It is designed to evaluate whether reasoning methods can identify and correct structured causal failures that arise during the reasoning… See the full description on the dataset page: https://huggingface.co/datasets/EdmondFU/Causal-Reasoning-Bench_CRBench.question-answering10K<n<100K5 likes253 downloads2mo agoHugging Face03botcoinmoney /domain-agnostic-causal-reasoning-tuning Domain-Agnostic Causal Reasoning Tuning Dataset Training data for fine-tuning language models on multi-hop document reasoning. Each example is a graded reasoning trace produced by a frontier AI agent solving a procedurally generated challenge from the Botcoin proof-of-inference network. The traces contain no real domain knowledge. Entities are fictional, numbers are random, and documents are generated deterministically from 128-bit seeds. The reasoning structure is what matters:… See the full description on the dataset page: https://huggingface.co/datasets/botcoinmoney/domain-agnostic-causal-reasoning-tuning.textquestion-answering10K<n<100K1 likes169 downloads6mo agoHugging Face04jmaasch /compositional_causal_reasoning –&nbsp;3k+ Hugging Face downloads – https://jmaasch.github.io/ccr/ Causal reasoning and compositional reasoning are two core aspirations in AI. Measuring these behaviors requires principled evaluation methods. Maasch et al. (2025) consider both behaviors simultaneously, under the umbrella of compositional causal reasoning (CCR): the ability to infer how causal measures compose and, equivalently, how causal quantities propagate through graphs. CCR.GB applies the… See the full description on the dataset page: https://huggingface.co/datasets/jmaasch/compositional_causal_reasoning.question-answering3 likes151 downloads1mo agoHugging Face05anonsubmission16 /CausalReasoningBenchmark Automated Causal Reasoning Benchmark Anonymized release for double-blind review. Author, affiliation, and prior-whitepaper material have been removed. The data, solutions, and evaluation pipeline are otherwise identical to the version under review. Overview The Automated Causal Reasoning Benchmark is a collection of real-world causal inference tasks drawn from 85 peer-reviewed research papers and three textbook-style collections. The benchmark contains 173 queries over… See the full description on the dataset page: https://huggingface.co/datasets/anonsubmission16/CausalReasoningBenchmark.tabularquestion-answeringn<1K0 likes40 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.