datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ShareGPT-Chinese-English-90k
ShareGPT-Chinese-English-90k Bilingual Human-Machine QA Dataset
A high-quality Chinese-English parallel bilingual human-machine QA dataset, covering user questions in real and complex scenarios. It is used for training high-quality dialogue models (more robust in instruction distribution than those datasets generated by repeatedly calling API interfaces to simulate machine-generated Q&A, like Moss)
Features:
Provides fully semantically equivalent Chinese-English parallel corpus… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/ShareGPT-Chinese-English-90k.Audio-Video-Engineering-Agentic-Tasks-1M
Audio/Video Engineering Agentic Tasks (1M)
Abstract
A highly specialized dataset comprising 1,029,459 in-context troubleshooting prompts and execution commands built for the deepest levels of media production. Unlike standard datasets that simulate clean, theoretical instructions, this matrix captures the chaotic, highly-detailed, and conversational reality of professional audio engineers, composers, and video editors mid-session. It is engineered to train multimodal AI… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/Audio-Video-Engineering-Agentic-Tasks-1M.TamilNadu-State-board-books-2025-Tamil-and-english-versions
Tamil Nadu School Textbooks — Tamil and English Structured Text
This dataset contains text extracted from 312 Tamil Nadu State Board school
textbooks for Standards 1–12. It covers Tamil- and English-medium books and
provides each retained book in two forms:
structured JSON with book metadata, ordered sections, typed content blocks,
source references, extraction statistics, and curation provenance;
Markdown for reading, inspection, and downstream text processing.
The source… See the full description on the dataset page: https://huggingface.co/datasets/Joemn/TamilNadu-State-board-books-2025-Tamil-and-english-versions.distill-gpt4-eng-chat
Description
Introducing dataset consisting of gpt4 answers to users requests. Queries were taken from allenai/WildChat-1M and causal-lm/instructions. Texts (requests and responses) were deleted in 3 cases:
either has non-english letters and special symbols
either has http-links
either has html blocks
either has perplexity more than 1.5*IQR + third quantile ( in some cases average perplexity value of sentences or maximum value was used )
EngineMT-QA
EngineMT-QA Dataset
Overview
EngineMT-QA is a large-scale, multi-task, multimodal dataset for Time-Series Question Answering (Time-Series QA). It enables research on aligning multivariate time-series signals with natural language through four key cognitive tasks:
Understanding
Perception
Reasoning
Decision-Making
The dataset is built on N-CMAPSS, simulating real-world aero-engine operational and maintenance scenarios. It supports the development and evaluation of… See the full description on the dataset page: https://huggingface.co/datasets/pandalin98/EngineMT-QA.agriculture-qa-english-only
Dataset Card for Dataset Name
This dataset contains question-answer pairs related to agriculture. The dataset can be used for tasks such as question answering, information retrieval, and natural language understanding in the agricultural domain. The questions cover various aspects of agriculture, including crop production, animal husbandry, soil management, and farming practices.
Dataset Details
he dataset is structured as a collection of JSON files, with each file… See the full description on the dataset page: https://huggingface.co/datasets/KisanVaani/agriculture-qa-english-only.Flutter-Code-with-Questions-Dataset-English
🧠 Flutter Code with Questions Dataset (English)
This repository contains a high-quality dataset of Flutter-related code snippets paired with automatically generated English technical questions. The dataset is intended for use in training and fine-tuning language models, coding assistants, and educational systems focused on Flutter development.
📂 Dataset Structure
The dataset is divided into 22 CSV files, each containing 200 entries. Every entry includes:
A… See the full description on the dataset page: https://huggingface.co/datasets/NoirZangetsu/Flutter-Code-with-Questions-Dataset-English.Electrical-engineering
To the electrical engineering community
This dataset contains Q&A prompts about electrical engineering, Kicad's EDA software features and scripting console Python codes.
Authors
STEM.AI: stem.ai.mtl@gmail.comWilliam Harbec
PolyDevTasks-Chinese_English_German
PolyDevTasks: 多语言软件开发智能任务
💻 Github 仓库
简体中文 | English | Deutsch
介绍
我们发布了 PolyDevTasks,这是一个包含超过 38 万条真实编码任务指令的数据集,涵盖 3 种自然语言(中文、英文、德语)和 8 种编程语言(C、C#、C++、Go、Java、JavaScript、Python、Rust)。不同于翻译或模板化的数据集,每条指令都是独立编写的,体现了特定语言与生态的习惯用法(例如 Go 的并发、C# 的 LINQ、UNIX I/O),并强调智能体的行为特征,如工具使用、网络操作、文件处理和优雅退出。PolyDevTasks 专为训练和评测 Agents 与 LLMs 在端到端软件工作流和跨语言泛化能力上的表现而设计。
数据统计
📂 文件级统计(每个 NL/PL 文件)
文件
数目
instruction均长
response均长
zh/c.jsonl
24,590
78.90
6,374.07… See the full description on the dataset page: https://huggingface.co/datasets/Mxode/PolyDevTasks-Chinese_English_German.engsaf
Engineering Short Answer Feedback
A collection of real short-answer responses from engineering exams across multiple engineering domains.
Background
In recent years, there has been a growing interest in using Artificial Intelligence (AI) to automate student assessment in education.
Among different types of assessments, summative assessments play a crucial role in evaluating a student's understanding level of a course.
Such examinations often involve short-answer… See the full description on the dataset page: https://huggingface.co/datasets/IsmaelMousa/engsaf.GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned
GLM-5.1-Reasoning-1M-Cleaned is a cleaned and reformatted derivative of Kassadin88/GLM-5.1-1000000x. It preserves the original four-subset layout (main, PHD-Science, Multilingual-STEM, Math) while converting every example into a unified SFT-ready schema with explicit conversations, input, output, domain, and meta fields.
This release was prepared from the original dataset published by Kassadin88.
Summary
Teacher model in the data: GLM-5.1… See the full description on the dataset page: https://huggingface.co/datasets/EngMuhammadAtef/GLM-5.1-Reasoning-1M-Cleaned.ekar_english
Dataset Card for ekar_english
Dataset Summary
New!(9/18/2022) E-KAR v1.1 is officially released (at the main branch), with a higher-quality English dataset! In v1.1, we further improve the Chinese-to-English translation quality of the English E-KAR, with over 600 problems and over 1,000 explanations manually adjusted. You can still find previous version (as in the paper) in the v1.0 branch in the repo. For more information please refer to… See the full description on the dataset page: https://huggingface.co/datasets/jiangjiechen/ekar_english.Scholarly-Epistemic-Engine
Dataset Card for Scholarly-Epistemic-Engine: arXiv cs.AI Corpus and Embeddings
This dataset contains the processed text, metadata, and semantic vector embeddings of approximately 90,000 scholarly articles from the arXiv Computer Science - Artificial Intelligence (cs.AI) category, spanning from 1993 to December 2024. It is designed to support Retrieval-Augmented Generation (RAG) systems and semantic knowledge discovery.
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/whyamanbhardwaj/Scholarly-Epistemic-Engine.WenYanWen_English_Parallel
Dataset Card for WenYanWen_English_Parallel
Dataset Summary
The WenYanWen_English_Parallel dataset is a multilingual parallel corpus in Classical Chinese (Wenyanwen), modern Chinese, and English. The Classical Chinese and modern Chinese parts are sourced from the NiuTrans/Classical-Modern dataset, while the corresponding English translations are generated using Gemini Pro.
Data Fields
info: A string representing the title or source information of the text.… See the full description on the dataset page: https://huggingface.co/datasets/KaifengGGG/WenYanWen_English_Parallel.unreal-engine-5.7-qa
Unreal Engine 5.7 Instruction-Tuning Dataset
Dataset Description
This dataset contains 122,199 high-quality, synthetic Question and Answer pairs specifically designed for instruction-tuning Large Language Models (LLMs) to become expert coding and architectural assistants for Unreal Engine 5.7.
Because Unreal Engine frequently deprecates older APIs (from UE4 to UE5) and introduces massive paradigm shifts (like Nanite, Lumen, and World Partition), standard… See the full description on the dataset page: https://huggingface.co/datasets/TunstallTensor/unreal-engine-5.7-qa.LLM_Electrical_Engineering_Educational_Synthetic_DialogDataset Card for LLM_Electrical_Engineering_Educational_Synthetic_Dialog
The full CJ Jones' synthetic dataset catalog is available at: https://datadeveloper1.gumroad.com
Want more? 🚀 Get the AI Startup Bundle from Gumroad.
Dataset Description
The LLM_Electrical_Engineering_Educational_Synthetic_Dialog dataset contains AI-generated conversational interactions designed for training large language models in electrical engineering education. This synthetic dialogue corpus simulates tutor-student… See the full description on the dataset page: https://huggingface.co/datasets/CJJones/LLM_Electrical_Engineering_Educational_Synthetic_Dialog.shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.taiwanese_english_translationThis new dataset is designed to solve this great NLP task and is crafted with a lot of care.ERI-Engineering-Reasoning-and-Instruction
ERI — Engineering Reasoning and Instruction
53,429 instruction–response pairs spanning nine engineering disciplines, seven
question types and three difficulty levels, for instruction-tuning and evaluating
language models on engineering reasoning.
from datasets import load_dataset
eri = load_dataset("mznaser/ERI-Engineering-Reasoning-and-Instruction")
civil = eri.filter(lambda r: r["field"] == "civil_engineering")
Schema
Every record has six fields:
Field… See the full description on the dataset page: https://huggingface.co/datasets/mznaser/ERI-Engineering-Reasoning-and-Instruction.Small-Life-Dataset-ru-eng
Russian-English Dialogue Dataset
🎯 Overview
A comprehensive bilingual dialogue dataset containing 50,000 high-quality question-answer pairs in Russian and English. The dataset is balanced across two main categories: programming/technical topics and general conversation.
Dataset Statistics:
📊 Total Dialogues: 50,000
🇷🇺 Russian: 25,135 (50.3%)
🇬🇧 English: 24,865 (49.7%)
💻 Coding Topics: 25,056 (50.1%)
💬 General Conversation: 24,944 (49.9%)
📑… See the full description on the dataset page: https://huggingface.co/datasets/SonexaAI/Small-Life-Dataset-ru-eng.nemotron-terminal-software_engineering
nemotron-terminal-software_engineering
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "software_engineering". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-software_engineering.Math_CoT_Arabic_English_Reasoning
Math CoT Arabic English Dataset
A high-quality, bilingual (English & Arabic) dataset for Chain-of-Thought (COT) reasoning in mathematics and related disciplines, developed by Miscovery AI.
Overview
Math-COT is a unique dataset designed to facilitate and benchmark the development of chain-of-thought reasoning capabilities in language models across mathematical domains. With meticulously crafted examples, explicit reasoning steps, and bilingual support, this dataset offers… See the full description on the dataset page: https://huggingface.co/datasets/miscovery/Math_CoT_Arabic_English_Reasoning.English-STEM-QA-MCQ-DatasetDataset Description:
This dataset is a large-scale collection of English STEM Question Answering (QA) data, containing 4,388,206 question-answer pairs, designed to support the development and training of advanced NLP systems and AI models for scientific understanding, reasoning, problem-solving, and educational learning in English.
The dataset consists of multiple-choice question answering (MCQA) samples across core STEM domains including Physics, Mathematics, Chemistry, Biology, and General… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English-STEM-QA-MCQ-Dataset.sdr-engineers-dataset
📡 Software-Defined Radio (SDR) for Engineers Dataset
A comprehensive, high-quality instruction-tuning, preference optimization (DPO), and RAG dataset compiled from the textbook "Software-Defined Radio for Engineers" (2018) by Travis F. Collins, Robin Getz, Di Pu, and Alexander M. Wyglinski (Artech House).
This dataset covers Software-Defined Radio (SDR) architecture, digital signal processing (DSP) fundamentals, probability in communications, digital modulations (PAM, QAM, PSK)… See the full description on the dataset page: https://huggingface.co/datasets/onkanat/sdr-engineers-dataset.ml-ai-engineer-sft
DuoNeural ML/AI Engineer SFT Dataset
A synthetic instruction-tuning dataset for training an LLM to be a useful pairing partner on ML/AI engineering work — debugging training runs, reasoning about architecture and infra choices, reviewing experiment design, and explaining core ML concepts with the specificity of someone who's actually run the experiments.
Why this dataset exists
Most general instruction-tuning data treats ML engineering questions the same as any… See the full description on the dataset page: https://huggingface.co/datasets/DuoNeural/ml-ai-engineer-sft.nuclear_eng_HW_dataset
Automated Grading of Handwritten STEM Homework: A Five-Stage Pipeline for Nuclear Engineering
📄 Read the full paper (PDF)
Contents
automatic_homework_grader_nuclear.pdf — Full technical report describing the five-stage pipeline
data/grading_records.parquet — Structured grading records for each student submission
data/results.parquet — Evaluation results and error taxonomy
LLM Grading of Handwritten STEM Homework (Nuclear Physics)
Per-question… See the full description on the dataset page: https://huggingface.co/datasets/mst-ai/nuclear_eng_HW_dataset.lilium_albanicum_eng_alb
Lilium Albanicum Eng-Alb
Task Categories:
Translation
Question-Answering
Conversational
Languages: English (en), Albanian (sq)
Size Categories: 100K < n < 1M
Dataset Card for "Lilium Albanicum"
Dataset Summary
The Lilium Albanicum dataset is a comprehensive English-Albanian and Albanian-English parallel corpus. The dataset includes original translations and extended synthetic Q&A pairs, which are designed to support and optimize LLM translation… See the full description on the dataset page: https://huggingface.co/datasets/noxneural/lilium_albanicum_eng_alb.e3c-crf-english
Dataset description
Here we realease the dataset to perform the Case Report Forms filling task obtained from The European Clinical Case Corpus as described in the paper Converting Annotated Clinical Cases into Structured Case Report Forms presented at the BioNLP workshop at ACL 2025.
The dataset is composed by a set patients with related clinical_note that describe their history and conditions. Each patient is uniquely identified by the document_id column.
The task consists of… See the full description on the dataset page: https://huggingface.co/datasets/NLP-FBK/e3c-crf-english.patent-engineering-diagrams-preview
Patent & Engineering Diagram Reasoning — Authentic Engineering Preview
NKO Data Labs · design-partner market probe · target commercial release: 25k+ figure-linked examples
A multimodal engineering-data concept linking technical figures to structured component/concept and description context for diagram understanding and technical reasoning.
Preview status
This repository now contains a small authentic source-backed engineering-diagram preview using selected U.S.… See the full description on the dataset page: https://huggingface.co/datasets/NKODATALABS/patent-engineering-diagrams-preview.DBNL-public-qa-english-translation
