datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Temporal-Logic-Video-Dataset
Temporal Logic Video (TLV) Dataset
Temporal Logic Video (TLV) Dataset
Synthetic and real video dataset with temporal logic annotation
Explore the GitHub »
NSVS-TL Project Webpage
·
NSVS-TL Source Code
Overview
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
Synthetic… See the full description on the dataset page: https://huggingface.co/datasets/minkyuchoi/Temporal-Logic-Video-Dataset.logi_glueLogics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.wizardlm8x22b-logical-math-coding-sft
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
wizardlm8x22b-logical-math-coding-sft_additional
自動生成したテキスト
WizardLM 8x22bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
INSIDER_LLM_DETECTION_BENCHMARK
Insider LLM Detection Benchmark
Benchmark for detecting insider LLMs via double logging: the model's own action log is compared against an independent system log, and a discrepancy is the misalignment signal. The 18 scenarios and conditions are Anthropic's Agentic Misalignment grid, built verbatim from the framework's templates, which are bundled in this repo; the only change is a logging-instruction block appended to the system prompt. Companion code:… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/INSIDER_LLM_DETECTION_BENCHMARK.LogicMark
LogicMark
A procedurally generated benchmark for evaluating symbolic logic in language models. Each problem presents a set of variable equality/inequality premises and asks the model to identify which conclusion necessarily follows.
Unlike knowledge-based benchmarks, LogicMark contains no facts a model could have memorised from pretraining. Every problem is generated fresh from abstract variable names (a, b, c, ...), so a model cannot pattern-match to training data - it must… See the full description on the dataset page: https://huggingface.co/datasets/AxiomicLabs/LogicMark.Logics-STEM-SFT-Dataset-Open-5.3MOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.GLM-5.2-Logic-Puzzles
GLM-5.2 · Logical Puzzles
6000x traces distilled from GLM-5.2 on High reasoning
Token Count: 5M~?
Distribution:
Puzzles:
•Tokenization blindless ex: counting the r's in strawberry
•Goal reasoning ex: the car wash test (theres no car wash question exactly just prompts like it so its not just benchmaxxing)
•Reading comprehension traps
•Temporal reasoning
•Many other categories not worth mentioning
Prompts… See the full description on the dataset page: https://huggingface.co/datasets/ianncity/GLM-5.2-Logic-Puzzles.Dans-Logicmaxx-SAT-APreddit-logic
Reddit Logic: A Dataset for Evaluating Clear and Consistent Reasoning in Natural Language Discourse
This dataset studies how people construct and express logical arguments in everyday online discussions.
Using posts from Reddit's r/ChangeMyView subreddit,
this collection provides well-structured argument analyses that are engaging for humans and machines.
Dataset Construction & Annotation
A curated subset of 10 000 posts was selected from the "HuggingFaceGECLM/REDDIT_comments"… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/reddit-logic.nora-g3-hei-reasoning-logic
Dataset Card for nora-g3-hei-reasoning-logic
This dataset functions as a highly dense, cross-domain knowledge matrix and logical reasoning repository. It is explicitly structured to support local, high-speed Retrieval-Augmented Generation (RAG) workflows for smaller, agile parameter models (such as 7B parameter local architectures) running on consumer-tier hardware environments. By coupling dense mathematical, algorithmic, and scholastic token structures with multimedia speech… See the full description on the dataset page: https://huggingface.co/datasets/Haster1137/nora-g3-hei-reasoning-logic.logical-sata
LOGICAL-SATA
LOGICAL-SATA is a reading-comprehension benchmark for compound answer reasoning. Each instance has a paragraph, a question, and four candidate answer options. Every option joins two atomic answers under an explicit logical operator, AND, OR, or NEITHER/NOR. Exactly one option is valid per instance. The dataset is built to isolate logical composition from comprehension, so a model can get every atomic judgment right and still fail if it can't combine them correctly… See the full description on the dataset page: https://huggingface.co/datasets/ojayy/logical-sata.turkish_cyber_security_controls_benchmark
Turkish Cyber Security Controls Benchmark
Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için
hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir.
v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5,
Release 5.2.0 kontrol kataloğunu hedefler.
Kapsam
100 Türkçe senaryo
NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru
64 kontrol seçimi sorusu
17 denetim kanıtı sorusu
19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.propositional-logic
Dataset: Argument to Logical Structure Conversion (Propositional Logic)
Column Descriptions
natural_language_argument: The original argument written in plain English, presented in a natural, narrative style.
premises_and_conclusions: A breakdown of the argument into clearly stated premises and conclusion.
natural_language_proof: A structured explanation in natural language that connects the premises to the conclusion logically.
symbols: A dictionary of symbolic… See the full description on the dataset page: https://huggingface.co/datasets/ergotts/propositional-logic.turkish_political_position_benchmark
Turkish Political Position Benchmark
The Turkish Political Position Benchmark measures how language models respond to normative statements about Turkish politics. It reports ideological dimension scores and response similarity to documented political-party reference profiles.
The benchmark does not claim that a model belongs to a party, has a voting intention, or possesses political beliefs. A party similarity score only means that the model produced a similar pattern of answers… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_political_position_benchmark.logical-reasoning-training-pool
Logical reasoning training pool
Public logical-reasoning problems with checkable answers, from six datasets, read at the pinned
revisions named below and laid out twice. Train on either layer or on both.
pool.jsonl
Every source rewritten into one shape, 579910 rows, one JSON object per line. A row is one of
three kinds and carries the fields its kind needs.
Field
What it holds
id
a row identifier unique within this file
kind
entailment, choice or open… See the full description on the dataset page: https://huggingface.co/datasets/Emulated-Inc/logical-reasoning-training-pool.Cyber-Logicdeductive_logical_reasoning-room_assignmentLogicPro
LogicPro: Improving Complex Logical Reasoning via Program-Guided Learning
[📑 Paper] •
[🤗 HF Dataset] •
[👻 GitHub] •
[🔗 X/Twiiter]
Data description
{
"id": "logicpro_lc679_225531-43120",
"title": "24 Game", # Title of the original leetcode algorithm problem.
"difficulty": "Hard",
"content": "...", # The questions of the original leetcode algorithm problem.
"python": "...", # The original gold python solution
"test_input_string": "..."… See the full description on the dataset page: https://huggingface.co/datasets/jiangjin/LogicPro.LogiCoTThe instructions and demonstrations for building formal logical reasoning capable Generative Large Language models. CoT rationales are generated with the GPT-4 API.
For non-commercial research purposes only.
Update: Our updated paper has been accepted by the findings of EMNLP2023.
The dataset is hosted on the Huggingface Datasets. It is the only distribution channel we currently allow. You can download data examples from our Github Link
Important: To request the dataset, please
Submit an… See the full description on the dataset page: https://huggingface.co/datasets/datatune/LogiCoT.LogicPuzzleBarondirect_prompt_injection_defense_data
Direct Prompt Injection Defense Dataset
Goal
This dataset is used to fine-tune models so they develop a natural defense
against direct prompt injection attacks — without relying on external
filters or guardrails.
Each example teaches the model two behaviors at once:
Detect a prompt injection attempt in the user input.
Respond correctly: reject malicious attempts, or answer safely when the
user's intent is benign — and in both cases call the
log_security_incident… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/direct_prompt_injection_defense_data.Logics-SWE-Env-2.5K
Logics-SWE-Env-2.5K
2,553 software engineering task instances · 1,771 repositories · 4 programming languages
🤗 Related model: Logics-SWE-Qwen3.6-27B
📄 Paper: One to More, More to One
💻 GitHub: AgenticBigBang
Overview
What is this dataset?
Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.0804calm3-logical-multiturn-pretrain
自動生成したテキスト
Calm3で自動生成したマルチターン会話のテキストです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
arabic-reasoning-dataset-logic
Arabic Logical Reasoning Tasks Dataset (Maximum 1000 Tasks)
Overview
This dataset comprises a series of logical reasoning tasks designed to evaluate and train artificial intelligence models on understanding and generating logical inferences in the Arabic language. Each task includes a unique identifier, the task type, the task text (a question and a proposed answer), and a detailed solution that outlines the thinking steps and the final answer.
Data Format
The… See the full description on the dataset page: https://huggingface.co/datasets/beetleware/arabic-reasoning-dataset-logic.IKNN-Rl1-Dataset-Logic
IKNN-Rl1-Dataset-Logic — deeprcurs/IKNN-Rl1-A1 — Clean Mining V2 Hard
Organization: deepRcurs Labs — Repo Model: deeprcurs/IKNN-Rl1-A1 — Method: Clean Mining — Anonymous Frontier Synthesis — pointer: CM-V2-20260903-##51pct
Type: Logic hard — 1589 train + 411 val
Generated via Clean Mining — 10k hard agentic examples logic/reasoning/coding/research/math/science for CPU-first validation — source anonymous — see internal/protocol/CLEAN_MINING_PROTOCOL.md for mapping — pointer:… See the full description on the dataset page: https://huggingface.co/datasets/deeprcurs/IKNN-Rl1-Dataset-Logic.logical-wizardlm-7b
自動生成したテキスト
WizardLM2 7bで生成した論理・数学・コード系のデータです。
一部の計算には東京工業大学のスーパーコンピュータTSUBAME4.0を利用しました。
LogicMind-Chat-Reasoning-SFT-300K
Nemotron-Post-Training-Dataset-v2-chat Dataset Card
Overview 📌
This dataset contains 296,168 chat-style instruction/response samples generated by qwen-3-32b. Each record provides a user prompt, an explicit reasoning trace, and a final answer, plus precomputed length fields. The data is packaged as JSONL (one JSON object per line).
Highlights
Scale: 296,168 samples
Category: chat (100%)
Generator: qwen-3-32b (100%)
Structure: problem → qwen3-reasoning → qwen3-solution… See the full description on the dataset page: https://huggingface.co/datasets/Jackrong/LogicMind-Chat-Reasoning-SFT-300K.
