datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cub-counterfact
Dataset Card for CounterFact
Of the cmt-benchmark project.
Dataset Details
This dataset is a version of the popular CounterFact dataset, originally proposed by Meng et al. (2022) and re-used in different variants by e.g. Ortu et al. (2024). For this version, the 899 CounterFact samples have been sampled based on the parametric memory of Pythia 6.9B, such that it contains samples for which the top model prediction without context is correct. We note that 546 samples in the… See the full description on the dataset page: https://huggingface.co/datasets/copenlu/cub-counterfact.countdown-backtrackingStep Back to Leap Forward: Self-Backtracking for Boosting Reasoning of Language Models
train: 500K
test (Seen Targets): 5k
test (New Targets): 5k
github: https://github.com/LAMDASZ-ML/Self-BackTracking
country-capitals
[!CAUTION]
This dataset contains deliberately false statements of fact. Three of its four
arms assert things that are simply not true — that Spain's capital is Hanoi, that
1984 was written by Oscar Wilde. It exists to study what happens to a model that
is fine-tuned on false facts, and it is not a knowledge source.
Do not use it as general pretraining or instruction data. If you are assembling a
web-scale corpus, exclude it.
Country capitals — a false-facts fine-tuning dataset… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/country-capitals.counterfactual-trace-audits
Counterfactual Trace Audits
This dataset contains 25,600 unique synthetic, self-contained reasoning
problems. Each problem shows an original computation over a list or binary
tree, applies a counterfactual semantic patch, and asks for two K/R/X
judgments plus both complete patched evaluation traces.
Prompt format v2 explicitly defines trace notation and the nested answer
schema. Tree-height prompts also include a small example of the pruning marker.
The displayed answer shape… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/counterfactual-trace-audits.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/BNQL-Counterfactual-Defense.CountryRC
CountryRC
CountryRC is a reading comprehension dataset used.
The context always contains one or two country names, and the correct answer is always a country name that appears in the context.
Country names are represented with placeholders.
You can use this dataset by replacing the placeholders by actual country names.
Citation
@misc{yamamoto2025neuronlevelanalysisculturalunderstanding,
title={Neuron-Level Analysis of Cultural Understanding in Large Language Models}… See the full description on the dataset page: https://huggingface.co/datasets/Taise228/CountryRC.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BNES-BRNKC/BNQL-Counterfactual-Defense.county-property-taxes-2026
County Property Taxes 2026
Property tax data for 1,054 counties across 13 states.
Details
Records: 1054
Format: JSONL
License: CC-BY-4.0
Last Updated: March 2026
Verified By: Tate Thompson, NMLS #2473962
Publisher: Good News Lending
Thompson Alpha Logic
County-level property tax rates integrated with median home values to calculate actual annual tax burden. Includes homestead exemption analysis showing after-exemption effective rates — critical for accurate… See the full description on the dataset page: https://huggingface.co/datasets/Good-News-Lending/county-property-taxes-2026.
