datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Logics-STEM-SFT-Dataset-Open-1.6M
Logics-STEM-SFT-Dataset-2.2M
📰 News
[2026.01.05]🔥 Release of our Techinical Report.
[2026.01.05]🔥 Release the first version of Logics-STEM-8B-SFT, Logics-STEM-8B-RL, /Logics-STEM-SFT-Dataset-Open-1.6M.
Overview
What is this dataset?
Logics-STEM-SFT-Dataset-2.2M is a curated long Chain-of-Thought (CoT) SFT dataset for STEM reasoning, built on top of high-quality open-source data and enhanced through a rigorous curation and distillation… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M.Logics-STEM-SFT-Dataset-Open-5.3MOmniParsingBench
🤗 Model | 📑 Technical Report | 💻 GitHub
OmniParsingBench is a comprehensive, large-scale, and high-quality evaluation corpus designed to rigorously evaluate the unified parsing capabilities of Multimodal Large Language Models (MLLMs) across diverse modalities.
Unlike traditional single-task benchmarks, OmniParsingBench assesses the full spectrum of parsing performance—from fundamental signal detection to complex semantic reasoning—across six primary domains: Document… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/OmniParsingBench.Logics-SWE-Env-2.5K
Logics-SWE-Env-2.5K
2,553 software engineering task instances · 1,771 repositories · 4 programming languages
🤗 Related model: Logics-SWE-Qwen3.6-27B
📄 Paper: One to More, More to One
Overview
What is this dataset?
Logics-SWE-Env-2.5K is a collection of repository-level software engineering tasks for research on coding agents and environment-based reinforcement learning. It contains 2,553 unique task instances from 1,771 GitHub repositories… See the full description on the dataset page: https://huggingface.co/datasets/Logics-MLLM/Logics-SWE-Env-2.5K.designer-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.LogicStack-LeetCodeextract from LogicStack-LeetCode
公众号「宫水三叶的刷题日记」刷穿 LeetCode 系列文章源码
包括 编程题目、解析、tag、题目url
根据 leetcode 原始题目网页,修正了一些 文件名 和 文件内容 中标注的难度不一致的文件样本
LogicSkills
LogicSkills — Training Data
Companion training releases for the LogicSkills
benchmark (EMNLP 2026 Findings). This Hugging Face dataset has three subsets:
symbolization, countermodel, and validity, each with a single train split.
Each subdirectory contains the corresponding data and task documentation.
Code for our data generation pipeline can be found here.
Dataset
Path
Size
Task
Symbolization
symbolization/
100k × 2 languages
sentence → FOL
Countermodel… See the full description on the dataset page: https://huggingface.co/datasets/mainlp/LogicSkills.
