datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Logical_Reasoning_Chainsecon_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.synthetic-indian-logical-reasoning-CoTyes
chinese-logic-sentiment-dataset
中文逻辑情感分析数据集 (由豆包API生成)
-->English
这是一个专门用于增强中文情感分析模型逻辑推理能力的数据集,由 豆包 API 生成并经过人工清洗/筛选。
它包含 反讽(Irony)、双重否定(Double Negative)、转折(Transition) 和 简单句(Simple) 四种逻辑类型。
数据结构
该数据集包含三个部分:
train.csv: 训练集,包含 2176 条样本。
val.csv: 验证集,包含 545 条样本。
test.csv: 测试集,包含 960 条样本。用于模型最终评估。
数据字段
Field
Description
text
中文文本内容
label
情感标签,0 表示负面,1 表示正
type
逻辑类型,包含反讽、双重否定、转折和简单句
sub_type
逻辑子类型,进一步细分逻辑结构
domain
领域 (影视/美食/旅游/生活/购物/社交)
使用指南
from… See the full description on the dataset page: https://huggingface.co/datasets/YiMeng-SYSU/chinese-logic-sentiment-dataset.logical-form-d87cef
logical-form-d87cef
Synthetic sensors test data: 50 rows in data.csv.
All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations.
Fields
sample_id: random identifier for this generated sample.
row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/Indigo-Patricia/logical-form-d87cef.gender_predictionCreative_Stories_Logical_ReasoningLogic-OA-SFTtrainlogicwaver-reasoning-v1
LogicWaver Adversarial Reasoning Benchmark
Adversarial Semantic Reasoning Benchmark — 20 handcrafted odd-one-out puzzles where SOTA LLMs fail but humans succeed. For LLM evaluation & Chain-of-Thought probing.
🚀 Live Demo: https://huggingface.co/spaces/Eviezonr08/logicwaver-demo
▶️ Try 20 puzzles interactively - no install!
Files
Reasoning_Puzzle_without_proline.csv — 20 puzzles for evaluation
with_proline/Reasoning_Puzzle_proline.csv — Same 20 + pro_line… See the full description on the dataset page: https://huggingface.co/datasets/Eviezonr08/logicwaver-reasoning-v1.prop_logicclinical-diagnostic-logic-fragility-atlas-v0.1What this dataset tests
Diagnostic reasoning as an unfolding narrative with branching choices.
The model must identifywhere the diagnostic manifold bifurcatesand how small inference errors amplifyinto different outcome basins.
Required outputs
critical logic junctures
irreversibility flags
branch entropy score
inference error map
amplification factor
outcome basin divergence report
harm gradient
recoverability index
ru-alpaca-logic
Это переработка в alpaca-friendly формат датасетов от:
MERA-evaluation[MERA]
Из датасета взяты и переработаны только subsets (lcs, parus, rcb, rummu, ruopenbookqa, rutie, ruworldtree)
Vikhrmodels[law_mc]
Датасет переработан с учетом неободимого формата.
Всего в train["input"] - input_ids: 20088 | Наибольшая длинна: 1801 | Количество overflow при ctx(1024): 18
Alice-Logic
Dataset Summary
This dataset contains 213 logic- and philosophy-inspired question/answer pairs framed within the persona of Alice — a whimsical yet rigorous reasoning guide.
Each entry consists of:
A fixed system prompt defining Alice’s character and response style.
A question (instruction or challenge) generated from a curated logic/philosophy curriculum.
An answer blending accurate logical reasoning with Wonderland-inspired metaphorical storytelling.
The dataset is designed to… See the full description on the dataset page: https://huggingface.co/datasets/D1rtyB1rd/Alice-Logic.Par-Four-Fineweb-Edu-Fortified-LogicFiltered for logic with the following scrip:
https://huggingface.co/datasets/Josephgflowers/Par-Four-Fineweb-Edu-Fortified-Logic/resolve/main/find-reason-fine.py
aristotel_antology_logic_md "и", "или", "но", "поэтому",
"также", "однако",
"если", "тогда"
clinical-diagnostic-logic-fragility-atlas-v0.2
Clinical Diagnostic Logic Fragility Atlas v0.2
What this is
A small dataset that tests one question:
Can you detect when diagnostic logic is moving toward fragility, not just carrying ambiguity?
This repo focuses on diagnostic logic breakdown under clinical reasoning pressure.
It models a system where:
diagnostic signal consistency may weaken
hypothesis conflict may rise
evidence may fragment
inference stability may erode before overt decision failure appears… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-diagnostic-logic-fragility-atlas-v0.2.Logic-1Nanbeige-3B-Logical-FailuresTechnical Challenge: Blind Spots of Nanbeige4.1-3B
Model Tested: Nanbeige/Nanbeige4.1-3B
Loading Protocol: The model was loaded via transformers on a standard Google Colab T4 GPU. To prevent CUDA Out-Of-Memory (OOM) errors, the weights were downcast using torch_dtype=torch.float16 and mapped to VRAM using device_map="auto".
Analysis of Blind Spots
As a base model lacking Supervised Fine-Tuning (SFT) or RLHF, Nanbeige4.1-3B exhibits severe zero-shot degradation. The 10 failures in this dataset… See the full description on the dataset page: https://huggingface.co/datasets/Liebert28/Nanbeige-3B-Logical-Failures.Arabic_Logicqalogic_types_full_desclogic_typessimplrops_rulebased_logicLogicLens
LogicLens
tags: reasoning, logical-evaluation, evidence-extraction
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'LogicLens' dataset comprises paragraphs of text that contain statements which require logical reasoning and evidence evaluation. Each paragraph is assessed for the presence of clear evidence supporting the claims made within. The dataset labels each entry with a binary 'evidence_present' indicator (1 for 'yes', 0… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/LogicLens.armada-logics-ft
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/ecaccam/armada-logics-ft.logic_testAI-Accelerator-RTL-Logic-2026This repository contains high-performance Verilog RTL logic for next-gen AI hardware (Matrix Engines, Systolic Arrays, and Quantization Blocks).
🛑 ACCESS & LICENSING
Access is currently Gated.
For Academic/Individual Research: Request access via Hugging Face.
For Commercial Licensing & Full Dataset: Please send a formal inquiry to:
📧 [mounikagandikota2003@gmail.com]
Note: Commercial use without a valid license is strictly prohibited
logical-textsqwen35-2b-logic-blind-spots
Qwen3.5-2B Logic Blind Spots
This small dataset contains examples of mistakes made by the base language model Qwen/Qwen3.5-2B-Base.
The goal was to test simple reasoning situations where the model should give a short and precise answer. In several cases the model produces incorrect outputs or does not follow the format requested in the prompt.
This dataset is not intended to be a formal benchmark. It is only an exploratory collection of examples that illustrate some blind spots of… See the full description on the dataset page: https://huggingface.co/datasets/EvertBuzon/qwen35-2b-logic-blind-spots.
