datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
prop_logic_puzzlelogicLogicGraph
LogicGraph: Benchmarking Multi-Path Logical Reasoning via Neuro-Symbolic Generation and Verification
📖 Overview
Evaluations of large language models (LLMs) primarily emphasize convergent logical reasoning, where success is defined by producing a single correct proof. However, many real-world reasoning problems admit multiple valid derivations, requiring models to explore diverse logical paths rather than committing to one route.
To address this limitation, we introduce… See the full description on the dataset page: https://huggingface.co/datasets/kkkarry/LogicGraph.Temporal-Logic-Video-Dataset
Temporal Logic Video (TLV) Dataset
Temporal Logic Video (TLV) Dataset
Synthetic and real video dataset with temporal logic annotation
Explore the GitHub »
NSVS-TL Project Webpage
·
NSVS-TL Source Code
Overview
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.logi_gluelsat_logic_games-analytical_reasoningNovel annotated evaluation dataset of LSAT logic games associated with paper:
Lost in the Logic: An Evaluation of Large Language Models’ Reasoning Capabilities on LSAT Logic Games
Arxiv: http://arxiv.org/pdf/2409.19012
If you find this dataset useful, please cite the paper!
@misc{malik2024lostlogicevaluationlarge,
title={Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games},
author={Saumya Malik},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saumyamalik/lsat_logic_games-analytical_reasoning.severity_ablation_logicGradients_Gradients_and_Text_Full_Logic_CaptionsLogicVistaLogicNLI
Dataset Card for "LogicNLI"
@inproceedings{tian-etal-2021-diagnosing,
title = "Diagnosing the First-Order Logical Reasoning Ability Through {L}ogic{NLI}",
author = "Tian, Jidong and
Li, Yitian and
Chen, Wenqing and
Xiao, Liqiang and
He, Hao and
Jin, Yaohui",
editor = "Moens, Marie-Francine and
Huang, Xuanjing and
Specia, Lucia and
Yih, Scott Wen-tau",
booktitle = "Proceedings of the 2021 Conference on Empirical… See the full description on the dataset page: https://huggingface.co/datasets/tasksource/LogicNLI.Chinese-Logic-Multiple-Choicelogical-entailmenthttps://github.com/google-deepmind/logical-entailment-dataset
@inproceedings{
evans2018can,
title={Can Neural Networks Understand Logical Entailment?},
author={Richard Evans and David Saxton and David Amos and Pushmeet Kohli and Edward Grefenstette},
booktitle={International Conference on Learning Representations},
year={2018},
url={https://openreview.net/forum?id=SkZxCk-0Z},
}
logical-fallacyhttps://github.com/causalNLP/logical-fallacy
@article{jin2022logical,
title={Logical fallacy detection},
author={Jin, Zhijing and Lalwani, Abhinav and Vaidhya, Tejas and Shen, Xiaoyu and Ding, Yiwen and Lyu, Zhiheng and Sachan, Mrinmaya and Mihalcea, Rada and Sch{\"o}lkopf, Bernhard},
journal={arXiv preprint arXiv:2202.13758},
year={2022}
}
SWE-Star
SWE-Star
Introduction
SWE-Star is a family of language models based on the Qwen2.5-Coder family and trained on the SWE-Star dataset. The dataset contains approximately 250k agentic coding trajectories distilled from Devstral-2-Small using SWE-Smith tasks.
The complete data generation, training, and evaluation pipeline is openly available in our GitHub repository, enabling anyone to reproduce our results.
Additional details are available in our blog posts.… See the full description on the dataset page: https://huggingface.co/datasets/LogicStar/SWE-Star.Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.libero-logic
LIBERO Logical State and Action Trajectories
This repository contains LIBERO robot manipulation trajectories augmented with per-frame logical state and logical action annotations.
The data is stored as HDF5 files under datasets/. Each suite has one directory, and each task has one HDF5 file containing multiple demonstrations.
Recommended Hugging Face Layout
Keep the repository organized like this:
.
├── README.md
├── requirements.txt
├── visualize_dataset.py
└──… See the full description on the dataset page: https://huggingface.co/datasets/Hoshipu/libero-logic.whittle-stop-kd
Correction - 27 August 2026
kd_mt_top32.npz is misaligned and must not be used. Its per-turn spans were
computed against a throwaway per-turn sequence and then stored against the full
conversation, so only 43 of 258 weight-8.0 positions land on the turn
terminator; the other 215 land on token 236, a partial UTF-8 byte.
kd_top32.npz is correctly aligned, but its tail weighting does not do what
the section below claims. The capture stops one position short of the
terminator, so… See the full description on the dataset page: https://huggingface.co/datasets/logic65/whittle-stop-kd.multi-zebra-logic
Dataset Card for the MultiZebraLogic dataset
This dataset includes zebra puzzles in 39 European and 5 non-European languages and in two sizes: 2x3 and 4x5. It can be used for evaluating logical reasoning ability.
The data has been generated using the code in this repo.
Dataset Details
Dataset Description
Zebra puzzles are a type of constraint satisfaction problem. They describe a number of objects, N_objects, that each have attributes… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/multi-zebra-logic.logical-reasoningLogical-Reasoning-1500-DataTroLL-Logic-Locking-based-Hardware-Trojanswizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
SWE-Smith
A extended version of the original SWE-smith-py dataset with more problem descriptions!
ppt-logic_termVideo-MME-v2-logic-only-replacedwizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
INSIDER_LLM_DETECTION_BENCHMARK
Insider LLM Detection Benchmark
Benchmark for detecting insider LLMs via double logging: the model's own action log is compared against an independent system log, and a discrepancy is the misalignment signal. The 18 scenarios and conditions are Anthropic's Agentic Misalignment grid, built verbatim from the framework's templates, which are bundled in this repo; the only change is a logging-instruction block appended to the system prompt. Companion code:… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/INSIDER_LLM_DETECTION_BENCHMARK.SPY
SPY: Enhancing Privacy with Synthetic PII Detection Dataset
We proudly present the SPY Dataset, a novel synthetic dataset for the task of Personal Identifiable Information (PII) detection. This dataset highlights the importance of safeguarding PII in modern data processing and serves as a benchmark for advancing privacy-preserving technologies.
Key Highlights
Innovative Generation: We present a methodology for developing the SPY dataset and compare it to other… See the full description on the dataset page: https://huggingface.co/datasets/mks-logic/SPY.LogicBench-v1.0
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Recently developed large language models (LLMs) have been shown to perform remarkably well on a wide range of language understanding tasks. But, can they really "reason" over the natural language? This question has been receiving significant research attention and many reasoning skills such as commonsense, numerical, and qualitative have been studied. However, the crucial skill pertaining… See the full description on the dataset page: https://huggingface.co/datasets/cogint/LogicBench-v1.0.first_rag_db_manuel_config_trial
Atlas Hospital Türkçe Medikal RAG Deneyi
Bu depo, bir metni parçalama, parçaları gömme (embedding), ChromaDB'ye kaydetme ve benzerlik eşiğiyle cevaplanabilirlik kararı verme adımlarını uçtan uca göstermek için hazırlanmış bir ödev çalışmasıdır.
Kaynak veri, umutertugrul/turkish-hospital-medical-articles veri setindeki Atlas Hospital bölümüdür. Ham dosyada 130 makale bulunur; metne göre yinelenen iki kayıt çıkarıldığında 128 benzersiz makale işlenir.
Bu çalışma eğitim amaçlıdır.… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/first_rag_db_manuel_config_trial.
