datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Temporal-Logic-Video-Dataset
Temporal Logic Video (TLV) Dataset
Temporal Logic Video (TLV) Dataset
Synthetic and real video dataset with temporal logic annotation
Explore the GitHub »
NSVS-TL Project Webpage
·
NSVS-TL Source Code
Overview
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.LogicBench-v1.0
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really "reason" over the natural language? This question has been receiving significant research attention and many reasoning skills such as commonsense, numerical, and qualitative have been studied. However, the crucial skill pertaining… See the full description on the dataset page: https://huggingface.co/datasets/cogint/LogicBench-v1.0.LogicInference_OA
Dataset Card for "LogicInference_OA"
This is an re-produce of the dataset from LogicInference Dataset in paper: https://openreview.net/pdf?id=HAGeIS_Lcg9.
The github page of LogicInference Dataset: https://github.com/google-research/google-research/tree/master/logic_inference_dataset.
This dataset is aimed to offer more dataset for Open Assistant project, depending on their demands, there three columns: INSTRUCTION, RESPONSE, SOURCE.
The results in this dataset is a little different… See the full description on the dataset page: https://huggingface.co/datasets/KK04/LogicInference_OA.GLM-5.2-Logic-Puzzles
GLM-5.2 · Logical Puzzles
6000x traces distilled from GLM-5.2 on High reasoning
Token Count: 5M~?
Distribution:
Puzzles:
•Tokenization blindless ex: counting the r's in strawberry
•Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing)
•Reading comprehension traps
•Temporal reasoning
•Many other categories not worth mentioning
Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.logical-sata
LOGICAL-SATA
LOGICAL-SATA is a reading-comprehension benchmark for compound answer reasoning. Each instance has a paragraph, a question, and four candidate answer options. Every option joins two atomic answers under an explicit logical operator, AND, OR, or NEITHER/NOR. Exactly one option is valid per instance. The dataset is built to isolate logical composition from comprehension, so a model can get every atomic judgment right and still fail if it can't combine them correctly… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-sata.logic-grid-puzzles-training-pool
Logic grid puzzles training pool
Logic grid puzzles: a row of positions, a handful of attributes with one value per position, and a
list of clues that together admit exactly one arrangement. Two sets drawn for this pool by
generators run here under the seeds recorded below, and two public datasets read at the pinned
revisions named below, laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 390945 rows, one JSON… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logic-grid-puzzles-training-pool.turkish_cyber_security_controls_benchmark
Turkish Cyber Security Controls Benchmark
Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için
hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir.
v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5,
Release 5.2.0 kontrol kataloğunu hedefler.
Kapsam
100 Türkçe senaryo
NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru
64 kontrol seçimi sorusu
17 denetim kanıtı sorusu
19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.econ_logic_qa
EconLogicQA
EconLogicQA is a benchmark designed to test the sequential reasoning skills of large language models (LLMs) in economics, business,
and supply chain management. It diverges from typical benchmarks by requiring models to understand and sequence multiple interconnected
events, capturing complex economic logics. The benchmark includes multi-event scenarios and a thorough suite of evaluations to assess
proficiency in economic contexts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/yinzhu-quan/econ_logic_qa.logical-reasoning-training-pool
Logical reasoning training pool
Public logical-reasoning problems with checkable answers, from two datasets whose licences allow
commercial use, read at the pinned revisions named below and laid out twice. Every problem is an
entailment problem: a block of premises, one conclusion, and whether the premises make the
conclusion true, false or neither. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 340951 rows, one JSON object per… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logical-reasoning-training-pool.logical-csqa
LOGICAL-COMMONSENSEQA
LOGICAL-COMMONSENSEQA reframes commonsense reasoning as logical composition over pairs of atomic statements using plausibility-level operators, AND, OR, and NEITHER/NOR. Each instance has a commonsense question and four candidate answer options, and every option joins two atomic answers under one of these operators. Exactly one option is valid per instance.
Most commonsense benchmarks rely on single-label evaluation, so it's never clear whether a model… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-csqa.deepseek-r1-autonomous-math-logic-cot-2026
📐 Enterprise DeepSeek-R1 Autonomous Mathematical & Logic CoT SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step hypothesis exploration, error discovery, and dynamic backtracking Chain-of-Thought (<thought>) reasoning trees for fine-tuning LLMs (DeepSeek-R1-Distill-Qwen, Qwen-2.5-Math, Llama-3.3, Mistral) into World-Class Olympiad Mathematicians and Formal Verification Agents.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/deepseek-r1-autonomous-math-logic-cot-2026.logic-problems-reasoning-dataset
Dataset Card for my-distiset-a26cd729
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/sdiazlor/my-distiset-a26cd729/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/sdiazlor/logic-problems-reasoning-dataset.gliclass-v3-logic-dataset
GLiClass‑V3 Logic Dataset
Rows 7 776 | Split train only | Format Parquet | Language EN | License Apache‑2.0
What it is
A length‑balanced corpus of single‑sentence prompts built purely for inducing reasoning in language models.
Why it helps
Teaches symbolic‑logic patterns and multi‑label behaviour.
Buckets cover 15 word‑length ranges (4 → 1,024) in equal proportions, exposing models to both tiny and very long inputs.
Each example has 1‑50 true and… See the full description on the dataset page: https://huggingface.co/datasets/knowledgator/gliclass-v3-logic-dataset.LogicMind-Chat-Reasoning-SFT-300K
Nemotron-Post-Training-Dataset-v2-chat Dataset Card
Overview 📌
This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line).
Highlights
Scale: 296,168 samples
Category: chat (100%)
Generator: qwen-3-32b (100%)
Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.LLMEval-Logic
LLMEval-Logic — Public 80% Release
A Solver-Verified Chinese Benchmark for Logical Reasoning of LLMs with Adversarial Hardening.
📄 Paper (arXiv): https://arxiv.org/abs/2605.19597
🌐 Project: https://llmeval.com/
🐙 Code & evaluation pipeline: https://github.com/llmeval/LLMEval-Logic
🤗 Dataset (this card): https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic
⚠️ This is the 80% public release
LLMEval-Logic was built through a three-stage audit pipeline: (a)… See the full description on the dataset page: https://huggingface.co/datasets/llmeval-fdu/LLMEval-Logic.grammar_logic_rhetoric_and_mathLOGIC-701-instruct
LOGIC-701 (instruct)
Based on https://huggingface.co/datasets/hivaze/LOGIC-701
Sources https://github.com/EvilFreelancer/LOGIC-701-instruct
LogicIFEval
LogicIFEval
For evaluation scripts, please refer to our GitHub repository: https://github.com/mianzhang/LogicIF
The dataset contains two splits:
full: Complete benchmark dataset (3,050 instructions)
mini: Mini version for quick evaluation (749 instructions)
Each line in the JSONL files contains a single evaluation example with the following structure:
{
"task_id": "string", // Unique identifier for the problem
"test_case_id": "int", // Test case number for… See the full description on the dataset page: https://huggingface.co/datasets/billmianz/LogicIFEval.Aristole
Dataset Name
This dataset contains instructions, inputs, and responses formatted for training language models. It is designed to help models understand and generate responses based on given instructions and inputs.
Dataset Structure
The dataset is structured with the following features:
Instruction: A string containing the task description or question.
Input: A string providing additional context or options.
Response: A string with the expected answer or completion.
designer-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.logicwaver-reasoning-v1
LogicWaver Adversarial Reasoning Benchmark
Adversarial Semantic Reasoning Benchmark — 20 handcrafted odd-one-out puzzles where SOTA LLMs fail but humans succeed. For LLM evaluation & Chain-of-Thought probing.
🚀 Live Demo: https://huggingface.co/spaces/Eviezonr08/logicwaver-demo
▶️ Try 20 puzzles interactively - no install!
Files
Reasoning_Puzzle_without_proline.csv — 20 puzzles for evaluation
with_proline/Reasoning_Puzzle_proline.csv — Same 20 + pro_line… See the full description on the dataset page: https://huggingface.co/datasets/Eviezonr08/logicwaver-reasoning-v1.ru-instruct-KAN-logic-v1
Russian Instruct KAN-Logic Dataset (v1)
Overview
ru-instruct-KAN-logic-v1 — это специализированный набор данных для instruction tuning (дообучения) языковых моделей на русском языке.
Основной фокус датасета — сложные логические рассуждения (Reasoning), математическое обоснование нейросетевых архитектур нового поколения (KAN - Kolmogorov-Arnold Networks) и теория распределенных вычислений.
Датасет содержит синтетические и курируемые пары instruction - output… See the full description on the dataset page: https://huggingface.co/datasets/K-Net-Labs/ru-instruct-KAN-logic-v1.Logic-ORiented-Retrieve
Logic-ORiented Retriever Enhancement Dataset
Dataset Description
This dataset is designed for training and evaluating Logic-ORiented Retriever Enhancement (LORE) models.
The dataset implements a three-tier contrastive learning framework with fine-grained sample classification:
P (Positive, label=1): Chunks sufficient to answer the query
N1 (Distractor, label=-1): Chunks used by LLM in query rewriting, seemingly relevant but unhelpful
N2 (Negative, label=0): Other… See the full description on the dataset page: https://huggingface.co/datasets/XiaSheng/Logic-ORiented-Retrieve.guru-logic-verl
GURU Logic VERL Dataset
Dataset Overview
This Hugging Face dataset contains 1,742 samples of logic reasoning problems from the GURU-RL-92k collection, specifically the logic and simulation splits after schema transformation. The data follows VERL (VerL format) specifications for reinforcement learning applications in logic reasoning tasks.
Key Features
Multi-domain Logic Reasoning: Covers ordering puzzles, zebra puzzles, graph problems, and ARC-AGI tasks… See the full description on the dataset page: https://huggingface.co/datasets/sungyub/guru-logic-verl.logic-rl-24k
Logic RL 24K
Logic RL 24K is a 24,461-example English reasoning mixture prepared for reinforcement learning with verifiable rewards (RLVR). It combines procedurally generated, algorithmically verifiable tasks from NVIDIA's Nemotron RL Reasoning Gym release with logic and table-reasoning tasks selected from LLM360's Guru RL 92K collection.
Every record contains one user message, a reference answer, and task-specific metadata that can be used to route the example to the… See the full description on the dataset page: https://huggingface.co/datasets/wflying/logic-rl-24k.logicLMLogic-ORiented-Test
XiaSheng/Logic-ORiented-Test
Logic-ORiented Test Dataset - Modified Test
This dataset contains modified test data for three different tasks:
hotpotqa_modified_test: Modified HotpotQA test questions
msmarco_modified_test: Modified MS MARCO test questions
musique_modified_test: Modified MuSiQue test questions
Each split contains questions that have been modified to test logic-oriented retrieval capabilities.
Dataset Structure
The dataset has three splits:… See the full description on the dataset page: https://huggingface.co/datasets/XiaSheng/Logic-ORiented-Test.italian-logic-repair-sft-dataset
Italian Logic Repair SFT Dataset
Teacher-backed synthetic Italian-first dataset designed for supervised fine-tuning repair. It targets exact arithmetic, concise direct QA, executable Python functions, JSON-only output, constraint following, stop behavior, and reasoning final-answer-marker behavior. Teacher outputs are used as candidates, then validated, corrected, or rejected by deterministic checks.
Dataset Details
Field
Value
Repository… See the full description on the dataset page: https://huggingface.co/datasets/SerFabio89/italian-logic-repair-sft-dataset.tensor-logic-wikipedia
Tensor Logic Wikipedia Knowledge Base
A structured knowledge base extracted from Wikipedia, designed for hybrid neural-symbolic reasoning.
Dataset Description
This dataset contains:
403,059 facts in Datalog-style format
210,188 entity embeddings (128 dimensions) learned from relationship patterns
Extracted from 37,000+ Wikipedia articles (Vital Articles + random sample)
Files
File
Description
Size
facts_only.tl
Clean facts in Relation(Subject… See the full description on the dataset page: https://huggingface.co/datasets/zekebass/tensor-logic-wikipedia.logic-trainingset-symb-structed-reformatted
Logic Reasoning and Proof Verification Dataset
Dataset Description
A comprehensive dataset for training and evaluating logical reasoning capabilities in language models.
Each example contains propositional logic problems that require formal reasoning to verify hypotheses.
The dataset includes 10427 problems with the following distribution:
- PROVED: 4291 examples where the hypothesis can be proven from the facts
- DISPROVED: 4262 examples where the hypothesis can be… See the full description on the dataset page: https://huggingface.co/datasets/RedaAlami/logic-trainingset-symb-structed-reformatted.
