CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /amazon_counterfactual_en Amazon Counterfactual Statements This dataset is the en-ext split from SetFit/amazon_counterfactual. As the original test set is rather small (1333 examples), a different split was created with 50-50 for training & testing. The dataset is described in amazon-multilingual-counterfactual-dataset / Paper It contains statements from Amazon reviews about events that did not or cannot take place. text10K<n<100K0 likes2.6k downloads5y agoHugging Face02ISLAM-PO /arab-dialects-20-countries-3m Dataset evaluation: See EVALUATION.md for schema checks, indexing status, and quality limitations. Viewer note: default is a lightweight preview; select full to load the complete corpus. Current Hub Validation Status Repository claim: 3,000,000 records Dataset Server indexed rows: 1,183,361 Dataset Server estimate: 2,064,964 The 3M target figure is a raw-repository claim and is not yet fully verified by the Hub index. Validate the JSONL files before publishing a definitive… See the full description on the dataset page: https://huggingface.co/datasets/ISLAM-PO/arab-dialects-20-countries-3m.texttext-generation1M<n<10M0 likes678 downloads1d agoHugging Face03copenlu /cub-counterfact Dataset Card for CounterFact Of the cmt-benchmark project. Dataset Details This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.tabularquestion-answering10K<n<100K0 likes317 downloads1y agoHugging Face04CG-Bench /CG-AV-Countinggated CG-AV-Counting Updates [2025/07/22] Since errors in a few clue annotations when converting frame indexes to timestamps, there were errors in the previous benchmark leaderboard, we have reevaluated all models and have updated the new leaderboard. Summary Despite progress in video understanding, current MLLMs struggle with counting tasks. Existing benchmarks are limited by short videos, close-set queries, lack of clue annotations, and weak… See the full description on the dataset page: https://huggingface.co/datasets/CG-Bench/CG-AV-Counting.textvisual-question-answering1K<n<10K5 likes250 downloads1y agoHugging Face05yangxw /countdown-backtrackingStep Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models train: 500K test (Seen Targets): 5k test (New Targets): 5k github: https://github.com/LAMDASZ-ML/Self-BackTracking textquestion-answering100K<n<1M0 likes232 downloads2y agoHugging Face06false-facts-finetuning /country-capitals [!CAUTION] This dataset contains deliberately false statements of fact. Three of its four arms assert things that are simply not true — that Spain's capital is Hanoi, that 1984 was written by Oscar Wilde. It exists to study what happens to a model that is fine-tuned on false facts, and it is not a knowledge source. Do not use it as general pretraining or instruction data. If you are assembling a web-scale corpus, exclude it. Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.textquestion-answering10K<n<100K0 likes184 downloads18d agoHugging Face07PersonaBias /counterfactuals Persona Bias Counterfactuals This dataset contains counterfactual examples used for persona-bias circuit discovery and intervention experiments. Repository Layout Hugging Face dataset config = model Hugging Face dataset split = counterfactual strategy task and axis are columns, not separate dataset configs data/<model>/<strategy>.jsonl.gz manifest.jsonl Strategies Split Meaning original Full original counterfactual set derived from… See the full description on the dataset page: https://huggingface.co/datasets/PersonaBias/counterfactuals.texttext-classification100K<n<1M0 likes165 downloads3mo agoHugging Face08HanSolo9682 /Flickr30k-CounterfactualsPaper: https://arxiv.org/abs/2402.13254 Project page: https://countercurate.github.io/ Code: https://github.com/HanSolo9682/CounterCurate text100K<n<1M4 likes128 downloads3y agoHugging Face09Parallel-Reasoning /countdown_problemstabular100K<n<1M0 likes119 downloads1y agoHugging Face10Raidriar-Dai /executable-counterfactuals Introduction This repo contains all training and evaluation datasets used in "Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code". This work has been published in ICLR 2026. Arxiv Paper Github Repo (Work in Progress) Counterfactual reasoning, a hallmark of intelligence, consists of three steps: inferring latent variables from observations (abduction), constructing alternative situations (interventions), and predicting the outcomes of the alternatives… See the full description on the dataset page: https://huggingface.co/datasets/Raidriar-Dai/executable-counterfactuals.tabular1K<n<10K0 likes101 downloads1mo agoHugging Face11divelab /countdowntabular1M<n<10M0 likes75 downloads1y agoHugging Face12Video-R1 /DVD-countingExtract From DVD: A Diagnostic Dataset for Multi-step Reasoning in Video Grounded Dialogue text1K<n<10K0 likes68 downloads2y agoHugging Face13MartialTerran /Eval_Counting_Letters_in_WordsLetters in Words Evaluation Dataset "The strawberry question is pretty much the new Turing Test for future AI" BlakeSergin OP 3mo agohttps://www.reddit.com/r/singularity/comments/1enqk04/how_many_rs_in_strawberry_why_is_this_a_very/ This dataset .json provides a simple yet effective way to assess the basic letter-counting abilities of Large Language Models (LLMs). (Try it on the new SmolLM2 models.) It consists of a set of questions designed to evaluate an LLM's capacity for: Understanding… See the full description on the dataset page: https://huggingface.co/datasets/MartialTerran/Eval_Counting_Letters_in_Words.textn<1K0 likes62 downloads2y agoHugging Face14KazeJiang /Multi-CounterFact Multi-CounterFact 🔍 Overview Multi-CounterFact is a multilingual benchmark for cross-lingual knowledge editing in large language models. While preserving the original evaluation structure for reliability, generality, and locality, it extends the original CounterFact dataset (Meng et al., 2022) from English to five languages: English, German, French, Japanese and Chinese. Each data instance represents a single editable factual association and contains: one target factual… See the full description on the dataset page: https://huggingface.co/datasets/KazeJiang/Multi-CounterFact.tabular100K<n<1M0 likes49 downloads6mo agoHugging Face15qualcomm /qualcomm-interactive-cooking-dataset-counterfactual-mistakes Qualcomm Interactive Cooking Dataset: Ego Counterfactual Mistakes Description This synthetic dataset contains mistake-intervention annotations for interactive cooking guidance. Each row contains video segment with instruction/feedback text pairs and their timestamps. Dataset Details Files: annotations.json Release statistics: Total rows: 25,087 Unique videos (dataset + video_id): 1,110 Rows by source dataset: CaptainCook4D: 4,969 Ego4D: 13,847 Ego-Exo4D: 6… See the full description on the dataset page: https://huggingface.co/datasets/qualcomm/qualcomm-interactive-cooking-dataset-counterfactual-mistakes.documenttext-generation10K<n<100K1 likes49 downloads5mo agoHugging Face16data-sci-project /ev-count-google-apitabular1K<n<10K0 likes48 downloads4d agoHugging Face17normalcomputing /wikiqa-counterfactualModel Card for Long-range Counterfactual WikiQA Github: https://github.com/normal-computing/extended-mind-transformers/ ArXiv: https://arxiv.org/abs/2406.02332 Original dataset by Abacus AI. Developed by: Normal Computing, Adapted from Abacus AI License: Apache 2.0 Long-range Counterfactual Retrieval Benchmark This benchmark is a modified wikiQA benchmark. The dataset is composed of Wikipedia articles (of 2-16 thousand tokens) and corresponding questions. We modify the… See the full description on the dataset page: https://huggingface.co/datasets/normalcomputing/wikiqa-counterfactual.textn<1K1 likes45 downloads2y agoHugging Face18NotoriousH2 /countdown-rlvr Countdown RLVR Qwen3-4B의 검증 가능한 추론 학습에 사용하는 Countdown 데이터셋입니다. 주어진 숫자를 각각 한 번 사용하여 목표값을 만드는 수식을 생성합니다. 1. 데이터 구성 분할 개수 숫자 개수 목표값 SHA-256 train 1,024 4 10~100 aa7abb6242d8ada72e55a6d8d0917e3618473ddf2f0f880b288814394b231131 validation 128 4 10~100 b06a1be3604d637aa19bd61af57aadf98fbbffcb8ef4db8d477fe8535c617497 test 256 4 10~100 416c02076321875cccfeed19f742e56048269b4b9d24112f6a2bee82ba301815 demo.jsonl에는 검증 흐름을 확인하는 숫자 3개 문제를 둡니다.… See the full description on the dataset page: https://huggingface.co/datasets/NotoriousH2/countdown-rlvr.tabulartext-generation1K<n<10K1 likes44 downloads1mo agoHugging Face19UKPLab /amazon_counterfactual_enAmazon Multilingual Counterfactual Dataset (https://arxiv.org/abs/2104.06893) The dataset contains sentences from Amazon customer reviews (sampled from Amazon product review dataset) annotated for counterfactual detection (CFD) binary classification. Counterfactual statements describe events that did not or cannot take place. Counterfactual statements may be identified as statements of the form – If p was true, then q would be true (i.e. assertions whose antecedent (p) and consequent (q) are… See the full description on the dataset page: https://huggingface.co/datasets/UKPLab/amazon_counterfactual_en.text10K<n<100K0 likes39 downloads4y agoHugging Face20electricsheepafrica /Constitution-of-all-54-African-countries Constitution of all 54 African countries | Africa (Electric Sheep Africa metadata inventory) Size category: n<1K - Formats: json - Sector: other_unclassified - Engineered by Electric Sheep Africa TL;DR This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context. What This Dataset Covers Public… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/Constitution-of-all-54-African-countries.texttabular-classificationn<1K0 likes38 downloads2mo agoHugging Face21MiaoMiaoYang /CXR-CounterFact CXR-CounterFact (CCF) Dataset We are pioneers in introducing counterfactual cause into reinforced custom-tuning of MLLMs, we are deeply aware of the scarcity of counterfactual CoT in downstream tasks, especially in the highly professional medical field. Thus, our aspiration is for the model to adeptly acclimate to the concept drift by itself, acquiring abundant knowledge with more and more data, but not exhibiting bias. In this context, a more realistic training dataset for… See the full description on the dataset page: https://huggingface.co/datasets/MiaoMiaoYang/CXR-CounterFact.textimage-to-text100K<n<1M3 likes37 downloads5mo agoHugging Face22xlr8harder /counterfactual-trace-audits Counterfactual Trace Audits This dataset contains 25,600 unique synthetic, self-contained reasoning problems. Each problem shows an original computation over a list or binary tree, applies a counterfactual semantic patch, and asks for two K/R/X judgments plus both complete patched evaluation traces. Prompt format v2 explicitly defines trace notation and the nested answer schema. Tree-height prompts also include a small example of the pruning marker. The displayed answer shape… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/counterfactual-trace-audits.tabularquestion-answering10K<n<100K0 likes33 downloads1mo agoHugging Face23BearNetworkChain /BNQL-Counterfactual-Defense 🚩 Γ Physics Engine — Canonical Definition Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆 最早提出時間:2025 年 6 月 19 日 原始來源:https://www.facebook.com/share/p/19cadcMTGo/ Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo 📌 0. 語義一致性設計層(Semantic Normalization Layer) 本文件定義 Γ Physics Engine 的標準語義行為規格,目的為: 在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。 📎 語義規則(強制一致) 為避免歧義,本文件採用以下規則: 中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/BNQL-Counterfactual-Defense.textquestion-answeringn<1K1 likes29 downloads4mo agoHugging Face24referencesource /iecc-climate-zone-by-county IECC/Building America climate zone by U.S. county, 2021 code cycle Canonical, always-current version: https://referencesource.org/iecc-climate-zone-by-county/ Machine-readable: https://referencesource.org/iecc-climate-zone-by-county/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-19 Stale after: 2028-08-18 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 3134 Which IECC climate zone (1-8, with… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/iecc-climate-zone-by-county.text1K<n<10K0 likes29 downloads1mo agoHugging Face25NousResearch /Letter-Counting-Evaltextn<1K5 likes28 downloads9mo agoHugging Face26es-heterogeneity /countdown-dataset ES Heterogeneity Countdown Countdown arithmetic data used for Evolution Strategies experiments under heterogeneous data allocation. Dataset splits train: approximately 3.79 million synthetically generated, solvable, and deduplicated Countdown problems. test: 2,000 held-out Countdown problems from the original evaluation set. Each example contains: id: example identifier numbers: input numbers that must each be used exactly once target: desired arithmetic result… See the full description on the dataset page: https://huggingface.co/datasets/es-heterogeneity/countdown-dataset.texttext-generation1M<n<10M0 likes28 downloads1mo agoHugging Face27twistshan /realistic-niah-count-mechanism-analysis Realistic NIAH count mechanism analysis Version 2 stores the paired geometry panel once. The default geometry_shared configuration contains 300 unique V4.4 stimulus rows: 200 discovery rows (seeds 1234-1253) and 100 held-out confirmation rows (seeds 1254-1263), with counts 1-10 balanced within every seed. Each pair_id is now one row rather than two duplicated mode rows. The common row contains the passage, gold records, slots, active needle spans, hard negatives, design metadata… See the full description on the dataset page: https://huggingface.co/datasets/twistshan/realistic-niah-count-mechanism-analysis.tabulartext-generationn<1K0 likes27 downloads1mo agoHugging Face28XiaoyanLi /medical_o1_sft_counter_cottext10K<n<100K2 likes25 downloads2y agoHugging Face29DataAttributionEval /Counterfact Overview This dataset is designed to evaluate data attribution methods for factual tracing. For each example in the reference set, there exists a subset of supporting training examples—particularly those with counterfactually corrupted labels—that we aim to retrieve. Importantly, all models are fine-tuned on the same training set, but each model has its own reference set, which captures the specific instances that expose counterfactual behavior during evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/DataAttributionEval/Counterfact.text10K<n<100K0 likes25 downloads1y agoHugging Face30HiTZ /counter-argument Dynamic Knowledge Integration for Evidence-Driven Counter-Argument Generation with Large Language Models 📋 Abstract This work investigates the role of dynamic external knowledge integration in improving counter-argument generation using Large Language Models (LLMs). While LLMs show promise in argumentative tasks, their tendency to generate lengthy, potentially unfactual responses highlights the need for more controlled and evidence-based approaches. We introduce a new… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/counter-argument.textn<1K1 likes25 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.