CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01medalpaca /medical_meadow_pubmed_causal Dataset Card for Pubmed Causal Dataset Summary This is the dataset used in the paper: Detecting Causal Language Use in Science Findings. Citation Information @inproceedings{yu-etal-2019-detecting, title = "Detecting Causal Language Use in Science Findings", author = "Yu, Bei and Li, Yingya and Wang, Jun", booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint… See the full description on the dataset page: https://huggingface.co/datasets/medalpaca/medical_meadow_pubmed_causal.textquestion-answering1K<n<10K10 likes946 downloads3y agoHugging Face02davidfoss /Synthetic-Causal-Reasoning-50k 🏭 Sovereign Synthetic Reasoning Dataset (400k) "High-Quality Chain-of-Thought Data at Scale." 📊 Overview This dataset contains 400,000 synthetic reasoning samples spanning 16 enterprise domains (Finance, Pharma, Legal, Cybersecurity, Supply Chain, etc.). It was generated using the Sovereign Generator, which produced 1.6 million samples and applied a strict quality filter (Top 25%) to retain only the most logically consistent and complex chains. Average Quality… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/Synthetic-Causal-Reasoning-50k.text100K<n<1M1 likes404 downloads9mo agoHugging Face03aryaman /causalgymCausalGym is a benchmark for comparing the performance of causal interpretability methods on a variety of simple linguistic tasks taken from the SyntaxGym evaluation set (Gauthier et al., 2020, Hu et al., 2020) and converted into a format suitable for interventional interpretability. The dataset includes train/dev/test splits (exactly as used in the experiments in the paper). The base/src columns are the prompts on which intervention is done. Each of these is a list of strings, with each… See the full description on the dataset page: https://huggingface.co/datasets/aryaman/causalgym.text10K<n<100K7 likes148 downloads3y agoHugging Face04ludwigw /causal-reasoning-benchmarks Causal Reasoning Benchmarks Datasets used in "On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning" (Deshmukh & Gupta, 2026). Dataset Structure train/transitivity_train.jsonl — 50,000 transitivity training examples train/dsep_train.jsonl — 50,000 d-separation training examples eval/length_eval.jsonl — 10,000 length generalization examples eval/branching_eval.jsonl — 10,000 branching structure examples eval/reversed_eval.jsonl — 10,000… See the full description on the dataset page: https://huggingface.co/datasets/ludwigw/causal-reasoning-benchmarks.texttext-classification100K<n<1M0 likes119 downloads5mo agoHugging Face05Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes104 downloads2mo agoHugging Face06Antix5 /vi-gym-causal-ascii Vi-Gym Causal ASCII Trajectories This dataset contains autoregressive trajectories of a Large Language Model (LLM) agent learning spatial reasoning and geometric drawing within a simulated Vi (Vim) editor environment. Warning This dataset is a direct derivation of the source material, it might therefore also contain content not suitable for all audiences. All authors of the original artwork have full ownership. Dataset Structure Each record is a discrete step… See the full description on the dataset page: https://huggingface.co/datasets/Antix5/vi-gym-causal-ascii.texttext-generation100K<n<1M0 likes103 downloads7mo agoHugging Face07CausalNLP /stride-lds STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations Ground-truth Linear Datamodeling Score (LDS) targets for the four nanochat pre-training models, plus the shared held-out test set. Each lds_<tag>.jsonl was produced by: sampling a pool of pre-training examples, drawing 256 random 30%-subsets, training a fresh nanochat from scratch on each subset, and recording per-example held-out test losses. The _meta header records the pool indices so scores defined over… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-lds.tabularothern<1K0 likes90 downloads3mo agoHugging Face08CausalLM /GPT-4-Self-Instruct-GermanHere we share a German dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly German. This dataset will be updated continuously. texttext-generation10K<n<100K30 likes87 downloads2y agoHugging Face09causaldrivebench /CausalDriveBench CausalDriveBench A benchmark for causal reasoning in autonomous driving built on top of nuScenes. Each sample bundles a curated causal scene graph, three flavours of multiple-choice / open-ended QA (active, dormant, distractor), and pointers to the raw nuScenes frames so the benchmark stays compact and license-clean. At a glance Subset uploaded: nuscenes Samples: 815 across 475 scenes Tasks: active QA, dormant QA, distractor QA, causal scene graphs Folder… See the full description on the dataset page: https://huggingface.co/datasets/causaldrivebench/CausalDriveBench.textquestion-answering1K<n<10K2 likes80 downloads5mo agoHugging Face10CausalLM /GPT-4-Self-Instruct-TurkishAs per the community's request, here we share a Turkish dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Turkish. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Turkish.text1K<n<10K24 likes75 downloads2y agoHugging Face11CausalLM /GPT-4-Self-Instruct-JapaneseHere we share a Japanese dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Japanese. This dataset will be updated continuously. text1K<n<10K18 likes66 downloads2y agoHugging Face12jifei0126 /CausalAirThe dataset is designed for fine-tuning a large language model for aviation accident analysis, including the reasoning process for identifying the causes of accidents. It consists of three fields: the narrative of the aviation accident, the reasoning process, and the accident cause. The dataset is divided into the SFT stage and the DPO stage. text10K<n<100K2 likes56 downloads9mo agoHugging Face13CausalLM /GPT-4-Self-Instruct-GreekAs per the community's request, here we share a Greek dataset synthesized using the OpenAI GPT-4 model with Self-Instruct, utilizing some excess Azure credits. Please feel free to use it. All questions and answers are newly generated by GPT-4, without specialized verification, only simple filtering and strict semantic similarity control have been applied. We hope that this will be helpful for fine-tuning open-source models for non-English languages, particularly Greek. This dataset will be… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/GPT-4-Self-Instruct-Greek.text1K<n<10K10 likes33 downloads2y agoHugging Face14crjojo /causal_kg_eval Logic-Aware Causal Knowledge Graph Dataset 이 데이터셋은 비구조화된 텍스트에서 추출된 논리적 관계와 인과 경로의 타당성을 평가하기 위해 설계되었습니다. 단순한 유사도 기반 검색을 넘어, 구조화된 지식을 활용한 고난도 추론 성능 측정을 목적으로 합니다. 1. 개요 (Overview) 목적: 정보 간의 선후 관계, 인과성 및 다단계(Multi-hop) 연결성 검증 데이터 형식: JSON (head, relation, tail) 핵심 기능: Semantic Noise 필터링 및 논리적 추론 경로(Reasoning Path) 제공 2. 관계 스키마 (Relation Schema) 본 데이터셋은 정보 간의 연결 강도와 성격에 따라 다음 4가지 관계를 정의합니다. Taxonomy (is_a): 상위 개념과 하위 개념 간의 계층적 분류 Causality (cause_of): 명확한 방향성을 가진… See the full description on the dataset page: https://huggingface.co/datasets/crjojo/causal_kg_eval.textquestion-answeringn<1K0 likes29 downloads5mo agoHugging Face15open-llm-leaderboard /CausalLM__14B-detailsgated Dataset Card for Evaluation run of CausalLM/14B Dataset automatically created during the evaluation run of model CausalLM/14B The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__14B-details.tabular10K<n<100K0 likes27 downloads2y agoHugging Face16open-llm-leaderboard /CausalLM__preview-1-hf-detailsgated Dataset Card for Evaluation run of CausalLM/preview-1-hf Dataset automatically created during the evaluation run of model CausalLM/preview-1-hf The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__preview-1-hf-details.tabular10K<n<100K0 likes23 downloads2y agoHugging Face17CausalLM /Refined-Anime-Textgated Refined Anime Text for Continual Pre-training of Language Models This is a subset of our novel synthetic dataset of anime-themed text, containing over 1M entries, ~440M GPT-4/3.5 tokens. This dataset has never been publicly released before. We are releasing this subset due to the community's interest in anime culture, which is underrepresented in general-purpose datasets, and the low quality of raw text due to the prevalence of internet slang and irrelevant content, making it… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Refined-Anime-Text.texttext-generation1M<n<10M274 likes22 downloads2y agoHugging Face18CausalLM /Retrieval-SFT-Chatgated Retrieval-Based Multi-Turn Chat SFT Synthetic Data A year ago, we released CausalLM/Refined-Anime-Text, a thematic subset of a dataset generated using the then state-of-the-art LLMs. This dataset comprises 1 million entries synthesized through long-context models that rewrote multi-document web text inputs, intended for continued pre-training. We are pleased to note that this data has been employed in various training scenarios and in studies concerning data and internet culture. In… See the full description on the dataset page: https://huggingface.co/datasets/CausalLM/Retrieval-SFT-Chat.textquestion-answering100K<n<1M61 likes20 downloads2y agoHugging Face19ARAYUN173 /arayun_173-system-law-symbolic-causal-coherence [DOI] https://doi.org/10.5281/zenodo.17186989 ARAYUN_173 – A System Law for Symbolic and Causal Coherence Corresponding author: ARAYUN_173 (Independent Research) E-mail: arayun173 [at] proton [dot] me Website: arayun173.com Date: September 2025 Audit Marker: SHA-256(ARAYUN_173|2025-09-04|Draft1) Contact: arayun173 [at] proton [dot] me ARAYUN_173 – A System Law for Symbolic and Causal Coherence Abstract ARAYUN_173 is not a concept but a system law. It establishes… See the full description on the dataset page: https://huggingface.co/datasets/ARAYUN173/arayun_173-system-law-symbolic-causal-coherence.textn<1K0 likes20 downloads3mo agoHugging Face20allenanie /causal_judgment Causal Judgment: Causal Reasoning with Moral, Intentional, and Counterfactual Analysis This task tests whether ultra-large language models are able to read a short story where multiple cause-and-effect events are introduced and answer causal questions such as "Did X cause Y?" in the same manner as humans would. Authors: Allen Nie (anie@cs.stanford.edu), Tobias Gerstenberg (gerstenberg@stanford.edu) Note: This repo is managed by the original author of this task. Please cite the… See the full description on the dataset page: https://huggingface.co/datasets/allenanie/causal_judgment.textn<1K0 likes17 downloads1y agoHugging Face21CausalNLP /stride-preproc-climbmix STRIDE: Preprocessed ClimbMix Tokenized ClimbMix sequences used for STRIDE's pretraining-data attribution experiments. There is one training pool per released nanochat depth (d12, d16, d20, d24) and one held-out test set shared across depths. These are the corpora the released nanochat operators attribute over, so STRIDE's column indices line up with the lines of these files. Files File Sequences Size Contents climbmix_train_d12.jsonl 1,317,003 3.8 GB training… See the full description on the dataset page: https://huggingface.co/datasets/CausalNLP/stride-preproc-climbmix.tabulartext-generation10M<n<100M0 likes17 downloads3mo agoHugging Face22open-llm-leaderboard /CausalLM__34b-beta-detailsgated Dataset Card for Evaluation run of CausalLM/34b-beta Dataset automatically created during the evaluation run of model CausalLM/34b-beta The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/CausalLM__34b-beta-details.tabular10K<n<100K0 likes16 downloads2y agoHugging Face23RexTRO111 /Causal-Conversationstextn<1K0 likes14 downloads2mo agoHugging Face24causal-seal /causal-seal-pilot Causal Seal — reference pilot An illustrative, self-contained snapshot showing the full chain a third party can follow without knowing the implementation that emitted it: exact commit SHA -> download / verify this snapshot -> open the STS trace (trace.jsonl) -> follow `causal_seal_ref` -> seal.json -> install / run the published verifier Files file role manifest.json entry point: what maps to what, sha256 per file. Git provides snapshot… See the full description on the dataset page: https://huggingface.co/datasets/causal-seal/causal-seal-pilot.textn<1K0 likes13 downloads1mo agoHugging Face25open-llm-leaderboard /DeepAutoAI__d2nwg_causal_gpt2_v1-detailsgated Dataset Card for Evaluation run of DeepAutoAI/d2nwg_causal_gpt2_v1 Dataset automatically created during the evaluation run of model DeepAutoAI/d2nwg_causal_gpt2_v1 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__d2nwg_causal_gpt2_v1-details.tabular10K<n<100K0 likes10 downloads2y agoHugging Face26REZ3LIET /NS-Causal-Inference-Synthetictext10K<n<100K0 likes9 downloads11mo agoHugging Face27mfmezger /deu-medical_meadow_pubmed_causaltext1K<n<10K0 likes7 downloads3y agoHugging Face28open-llm-leaderboard /DeepAutoAI__causal_gpt2-detailsgated Dataset Card for Evaluation run of DeepAutoAI/causal_gpt2 Dataset automatically created during the evaluation run of model DeepAutoAI/causal_gpt2 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/DeepAutoAI__causal_gpt2-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face29zuzannad1 /causal_datasets_mixedtext10K<n<100K0 likes6 downloads6mo agoHugging Face30CausaLab /causalab-graph-configs-review CausaLab Causal Graph Configuration Dataset Anonymous review release for NeurIPS 2026 Evaluations and Datasets. This package contains the synthetic causal graph configuration suites used by the active experiments in the submitted CausaLab paper. Each JSONL file contains 50 graph configurations, one JSON object per line. Contents . ├── data/ # 19 JSONL graph-suite files, 950 total records ├── manifest.json # file purposes, record… See the full description on the dataset page: https://huggingface.co/datasets/CausaLab/causalab-graph-configs-review.texttext-classificationn<1K0 likes6 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.