datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.cub-counterfact
Dataset Card for CounterFact
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.counterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/counterfactual-pendulum-multilingual.countdown-backtrackingStep Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
train: 500K
test (Seen Targets): 5k
test (New Targets): 5k
github: https://github.com/LAMDASZ-ML/Self-BackTracking
country-capitals
[!CAUTION]
This dataset contains deliberately false statements of fact. Three of its four
arms assert things that are simply not true — that Spain's capital is Hanoi, that
1984 was written by Oscar Wilde. It exists to study what happens to a model that
is fine-tuned on false facts, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are assembling a
web-scale corpus, exclude it.
Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.counterfactual-pendulum-multilingual
📌 Dataset Summary
When a Vision-Language Model (VLM) is given an image along with a text prompt containing contradictory or misleading information, how does it react? Does it rely on the visual evidence, succumb to textual bias, or honestly abstain when faced with unresolvable conflict?
This dataset adapts the Counterfactual Pendulum scenario across two visual conflict dimensions:
Angular (Angle): Conflict in the pendulum's angle of inclination.
Light: Conflict in the light… See the full description on the dataset page: https://huggingface.co/datasets/akanshjain37/counterfactual-pendulum-multilingual.counterfactual-hotpotqa
Counterfactual HotpotQA
A counterfactual yes/no QA dataset built from HotpotQA (distractor
setting) by minimally editing one answer-bearing supporting fact so the gold answer flips
(YES → NO or NO → YES), while keeping the question and entity names fixed.
Built for probing whether a Reading Comprehension / RAG model actually follows the evidence it is
given, rather than falling back on parametric/prior knowledge: for each example, a model is asked
the same question over the… See the full description on the dataset page: https://huggingface.co/datasets/prajaktakini/counterfactual-hotpotqa.countdown-es-grpo-0.1
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
This dataset contains prepared splits of the COUNTDOWN task for comparing Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) methods for LLM fine-tuning.
Dataset Statistics
Training split: 10.0% of available data
Training samples: 200
Validation samples: 1,800
Test samples: 200 (reserved for final evaluation)
Data Format
Each example contains:
data: The input… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-es-grpo-0.1.Country-city-animals
Country-city-animals: a dataset of synthetic facts, with corresponding corpora and reasoning tasks
Country-city-animals is a dataset of simple synthetic facts about countries, cities, and animals. The facts are provided in both triplet form and in text form, and can be used to train or finetune language models for studying knowledge learning from text. A variety of reasoning tasks are also provided to evaluate whether a model has learned the facts and can generalize them in… See the full description on the dataset page: https://huggingface.co/datasets/xiaozeroone/Country-city-animals.counterfactual-trace-audits
Counterfactual Trace Audits
This dataset contains 25,600 unique synthetic, self-contained reasoning
problems. Each problem shows an original computation over a list or binary
tree, applies a counterfactual semantic patch, and asks for two K/R/X
judgments plus both complete patched evaluation traces.
Prompt format v2 explicitly defines trace notation and the nested answer
schema. Tree-height prompts also include a small example of the pruning marker.
The displayed answer shape… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/counterfactual-trace-audits.countdown-full
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
Full Countdown dataset (2100 train + 100 test samples) for mathematical reasoning and arithmetic expression generation
Dataset Statistics
Training split: 100.0% of available data
Training samples: 2,100
Validation samples: 0
Test samples: 100 (reserved for final evaluation)
Data Format
Each example contains:
- data: The input prompt/question
- answer: Ground truth answer
-… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-full.Country-city-animals
Country-city-animals: a dataset of synthetic facts, with corresponding corpora and reasoning tasks
Country-city-animals is a dataset of simple synthetic facts about countries, cities, and animals. The facts are provided in both triplet form and in text form, and can be used to train or finetune language models for studying knowledge learning from text. A variety of reasoning tasks are also provided to evaluate whether a model has learned the facts and can generalize them in… See the full description on the dataset page: https://huggingface.co/datasets/amounts-tidings/Country-city-animals.countdown-es-grpo-0.4
COUNTDOWN Dataset for ES vs GRPO Comparison
Dataset Description
This dataset contains prepared splits of the COUNTDOWN task for comparing Evolution Strategies (ES) and Group Relative Policy Optimization (GRPO) methods for LLM fine-tuning.
Dataset Statistics
Training split: 40.0% of available data
Training samples: 800
Validation samples: 1,200
Test samples: 200 (reserved for final evaluation)
Data Format
Each example contains:
data: The input… See the full description on the dataset page: https://huggingface.co/datasets/alphaXiv/countdown-es-grpo-0.4.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/BNQL-Counterfactual-Defense.CountryRC
CountryRC
CountryRC is a reading comprehension dataset used.
The context always contains one or two country names, and the correct answer is always a country name that appears in the context.
Country names are represented with placeholders.
You can use this dataset by replacing the placeholders by actual country names.
Citation
@misc{yamamoto2025neuronlevelanalysisculturalunderstanding,
title={Neuron-Level Analysis of Cultural Understanding in Large Language Models}… See the full description on the dataset page: https://huggingface.co/datasets/Taise228/CountryRC.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/BNQL-Counterfactual-Defense.multilingual-counterfactual
Multilingual Counterfactual
A multilingual counterfactual MCQ dataset built from COCO-Counterfactual.
Each row contains an image, two captions (original vs counterfactual), and a multiple-choice question probing the difference between image content and text description.
Languages
Language
Code
Status
English
en
✅ Done
Hindi
hi
✅ Done
Urdu
ur
✅ Done
Telugu
te
✅ Done
Bahasa Indonesia
id
✅ Done
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/feliren/multilingual-counterfactual.leg-counting
leg-counting
Leg counting task - count total legs given a list of animals with quantities.
Dataset Structure
This dataset is in Hugging Face datasets format. Load it with:
from datasets import load_dataset
dataset = load_dataset("Tyrion279/leg-counting")
county-property-taxes-2026
County Property Taxes 2026
Property tax data for 1,054 counties across 13 states.
Details
Records: 1054
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Tate Thompson, NMLS #2473962
Publisher: Good News Lending
Thompson Alpha Logic
County-level property tax rates integrated with median home values to calculate actual annual tax burden. Includes homestead exemption analysis showing after-exemption effective rates — critical for accurate… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/county-property-taxes-2026.multilingual-counterfactual
Multilingual Counterfactual
A multilingual counterfactual MCQ dataset built from COCO-Counterfactual.
Each row contains an image, two captions (original vs counterfactual), and a multiple-choice question probing the difference between image content and text description.
Languages
Language
Code
Status
English
en
✅ Done
Hindi
hi
✅ Done
Urdu
ur
✅ Done
Telugu
te
✅ Done
Bahasa Indonesia
id
✅ Done
Columns
Column
Type… See the full description on the dataset page: https://huggingface.co/datasets/apart-global-south-hack/multilingual-counterfactual.
